JOURNAL ARTICLE

Learning Imbalanced Data with Vision Transformers

Abstract

The real-world data tends to be heavily imbalanced and severely skew the data-driven deep neural networks, which makes Long-Tailed Recognition (LTR) a massive challenging task. Existing LTR methods seldom train Vision Transformers (ViTs) with Long-Tailed (LT) data, while the off-the-shelf pretrain weight of ViTs always leads to unfair comparisons. In this paper, we systematically investigate the ViTs' performance in LTR and propose LiVT to train ViTs from scratch only with LT data. With the observation that ViTs suffer more severe LTR problems, we conduct Masked Generative Pretraining (MGP) to learn generalized features. With ample and solid evidence, we show that MGP is more robust than supervised manners. Although Binary Cross Entropy (BCE) loss performs well with ViTs, it struggles on the LTR tasks. We further propose the balanced BCE to ameliorate it with strong theoretical groundings. Specially, we derive the unbiased extension of Sigmoid and compensate extra logit margins for deploying it. Our Bal-BCE contributes to the quick convergence of ViTs in just a few epochs. Extensive experiments demonstrate that with MGP and Bal-BCE, LiVT successfully trains ViTs well without any additional data and outperforms comparable state-of-the-art methods significantly, e.g., our ViT-B achieves 81.0% Top-1 accuracy in iNaturalist 2018 without bells and whistles. Code is available at https://github.com/XuZhengzhuo/LiVT.

Keywords:
Computer science Artificial intelligence Train Artificial neural network Transformer Deep learning Machine learning Granularity Pattern recognition (psychology) Engineering

Metrics

46
Cited By
11.75
FWCI (Field Weighted Citation Impact)
114
Refs
0.98
Citation Normalized Percentile
Is in top 1%
Is in top 10%

Citation History

Topics

Domain Adaptation and Few-Shot Learning
Physical Sciences →  Computer Science →  Artificial Intelligence
Advanced Neural Network Applications
Physical Sciences →  Computer Science →  Computer Vision and Pattern Recognition
Multimodal Machine Learning Applications
Physical Sciences →  Computer Science →  Computer Vision and Pattern Recognition

Related Documents

JOURNAL ARTICLE

Investigating vision transformers for imbalanced ocular image classification with explainable ai

Syed Sohaib AliSuman Kumar Swarnkar

Journal:   Cuestiones de Fisioterapia Year: 2025 Vol: 54 (5)Pages: 1112-1131
JOURNAL ARTICLE

Evaluating Robustness of Vision Transformers on Imbalanced Datasets (Student Abstract)

Kevin LiRahul DuggalDuen Horng Chau

Journal:   Proceedings of the AAAI Conference on Artificial Intelligence Year: 2023 Vol: 37 (13)Pages: 16252-16253
BOOK-CHAPTER

Imbalanced Data Learning

Edward Yi Chang

Year: 2011 Pages: 191-211
JOURNAL ARTICLE

Federated Fuzzy Learning with Imbalanced Data

Lukas DustMarina López MurciaAndreas MakilaPetter NordinNing XiongFrancisco Herrera

Journal:   2021 20th IEEE International Conference on Machine Learning and Applications (ICMLA) Year: 2021 Pages: 1130-1137
© 2026 ScienceGate Book Chapters — All rights reserved.