Normalized Contrastive Learning for Text-Video Retrieval

Yookoon Park; Mahmoud Azab; Bo Xiong; Seungwhan Moon; Florian Metze; Gourab Kundu; Kirmani Ahmed

Conference ProceedingsOPEN ACCESS

Normalized Contrastive Learning for Text-Video Retrieval

Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022 (2022) 248-260

DOI: 10.18653/v1/2022.emnlp-main.17

5Citations

31Readers

Abstract

Cross-modal contrastive learning has led the recent advances in multimodal retrieval with its simplicity and effectiveness. In this work, however, we reveal that cross-modal contrastive learning suffers from incorrect normalization of the sum retrieval probabilities of each text or video instance. Specifically, we show that many test instances are either over- or under-represented during retrieval, significantly hurting the retrieval performance. To address this problem, we propose Normalized Contrastive Learning (NCL) which utilizes the Sinkhorn-Knopp algorithm to compute the instance-wise biases that properly normalize the sum retrieval probabilities of each instance so that every text and video instance is fairly represented during cross-modal retrieval. Empirical study shows that NCL brings consistent and significant gains in text-video retrieval on different model architectures, with new state-of-the-art multimodal retrieval metrics on the ActivityNet, MSVD, and MSR-VTT datasets without any architecture engineering.

References Powered by Scopus

View more at Scopus

Cited by Powered by Scopus

View more at Scopus

Cite

CITATION STYLE

APA

Park, Y., Azab, M., Xiong, B., Moon, S., Metze, F., Kundu, G., & Ahmed, K. (2022). Normalized Contrastive Learning for Text-Video Retrieval. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022 (pp. 248–260). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2022.emnlp-main.17

Readers' Seniority

PhD / Post grad / Masters / Doc 7

64%

Researcher 3

27%

Lecturer / Post doc 1

Readers' Discipline

Computer Science 12

80%

Medicine and Dentistry 1

Linguistics 1

Neuroscience 1

Normalized Contrastive Learning for Text-Video Retrieval

Abstract

References Powered by Scopus

Deep residual learning for image recognition

ImageNet: A Large-Scale Hierarchical Image Database

Unsupervised Feature Learning via Non-parametric Instance Discrimination

Cited by Powered by Scopus

Learning Visual Representations via Language-Guided Sampling

Learning to Ground Instructional Articles in Videos through Narrations

Register to see more suggestions

Cite

Readers' Seniority

Readers' Discipline