Improving Vision Transformers to Learn Small-Size Dataset From Scratch

This paper proposes various techniques that help Vision Transformer (ViT) to learn small-size datasets from scratch successfully. ViT, which applied the transformer structure to the image classification task, has outperformed convolutional neural networks, recently. However, the high performance of...

Full description

Bibliographic Details
Published in:IEEE Access
Main Authors: Seunghoon Lee, Seunghyun Lee, Byung Cheol Song
Format: Article
Language:English
Published: IEEE 2022-01-01
Subjects:
Online Access:https://ieeexplore.ieee.org/document/9957006/