Bag of Tricks for Long-Tailed Visual Recognition with Deep Convolutional Neural Networks

121Citations
Citations of this article
142Readers
Mendeley users who have this article in their library.

Abstract

In recent years, visual recognition on challenging long-tailed distributions, where classes often exhibit extremely imbalanced frequencies, has made great progress mostly based on various complex paradigms (e.g., meta learning). Apart from these complex methods, simple refinements on training procedure also make contributions. These refinements, also called tricks, are minor but effective, such as adjustments in the data distribution or loss functions. However, different tricks might conflict with each other. If users apply these long-tail related tricks inappropriately, it could cause worse recognition accuracy than expected. Unfortunately, there has not been a scientific guideline of these tricks in the literature. In this paper, we first collect existing tricks in long-tailed visual recognition and then perform extensive and systematic experiments, in order to give a detailed experimental guideline and obtain an effective combination of these tricks. Furthermore, we also propose a novel data augmentation approach based on class activation maps for long-tailed recognition, which can be friendly combined with re-sampling methods and shows excellent results. By assembling these tricks scientifically, we can outperform state-of-the-art methods on four long-tailed benchmark datasets, including ImageNet-LT and iNaturalist 2018. Our code is open-source and available at https://github.com/zhangyongshun/BagofTricks-LT.

References Powered by Scopus

Deep residual learning for image recognition

174328Citations
N/AReaders
Get full text

ImageNet: A Large-Scale Hierarchical Image Database

51115Citations
N/AReaders
Get full text

Microsoft COCO: Common objects in context

28859Citations
N/AReaders
Get full text

Cited by Powered by Scopus

A Survey on Long-Tailed Visual Recognition

94Citations
N/AReaders
Get full text

Nested Collaborative Learning for Long-Tailed Visual Recognition

79Citations
N/AReaders
Get full text

Long-tailed Visual Recognition via Gaussian Clouded Logit Adjustment

75Citations
N/AReaders
Get full text

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Cite

CITATION STYLE

APA

Zhang, Y., Wei, X. S., Zhou, B., & Wu, J. (2021). Bag of Tricks for Long-Tailed Visual Recognition with Deep Convolutional Neural Networks. In 35th AAAI Conference on Artificial Intelligence, AAAI 2021 (Vol. 4B, pp. 3447–3455). Association for the Advancement of Artificial Intelligence. https://doi.org/10.1609/aaai.v35i4.16458

Readers' Seniority

Tooltip

PhD / Post grad / Masters / Doc 53

91%

Researcher 4

7%

Professor / Associate Prof. 1

2%

Readers' Discipline

Tooltip

Computer Science 58

87%

Engineering 5

7%

Agricultural and Biological Sciences 2

3%

Economics, Econometrics and Finance 2

3%

Save time finding and organizing research with Mendeley

Sign up for free