Bibliographic record

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Authors: Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, Neil Houlsby
Publication year: 2020
OA status: unknown

Need access?

Ask circulation staff for physical copies or request digital delivery via Ask a Librarian.

Abstract

While the Transformer architecture has become the de-facto standard for natural language processing tasks, its applications to computer vision remain limited. In vision, attention is either applied in conjunction with convolutional networks, or used to replace certain components of convolutional networks while keeping their overall structure in place. We show that this reliance on CNNs is not necessary and a pure transformer applied directly to sequences of image patches can perform very well on image classification tasks. When pre-trained on large amounts of data and transferred to multiple mid-sized or small image recognition benchmarks (ImageNet, CIFAR-100, VTAB, etc.), Vision Transformer (ViT) attains excellent results compared to state-of-the-art convolutional networks while requiring substantially fewer computational resources to train.

Copies & availability

Realtime status across circulation, reserve, and Filipiniana sections.

Self-checkout (no login required)

Enter your student ID, system ID, or full name directly in the table.
Provide your identifier so we can match your patron record.
Choose Self-checkout to send the request; circulation staff are notified instantly.

Barcode	Location	Material type	Status	Action
No holdings recorded.

Digital files

Preview digitized copies when embargo permits.

No digital files uploaded yet.

Links & eResources

Access licensed or open resources connected to this record.

publisher Proxyable