2021
ViT: An Image is Worth 16x16 Words
Alexey Dosovitskiy, Lucas Beyer, +10
Vision Transformer (ViT) explained: how splitting images into 16x16 patches enables pure transformer architecture for state-of-the-art image recognition.
- Image Recognition
- Transformers
- Computer Vision
- +1 more tags
