Why are Sequence-to-Sequence Models So Dull? Understanding the Low-Diversity Problem of Chatbots

Shaojie Jiang; Maarten de Rijke

Conference ProceedingsOPEN ACCESS

Why are Sequence-to-Sequence Models So Dull? Understanding the Low-Diversity Problem of Chatbots

Proceedings of the 2018 EMNLP Workshop SCAI 2018: The 2nd International Workshop on Search-Oriented Conversational AI (2018) 81-86

DOI: 10.18653/v1/w18-5712

27Citations

125Readers

Abstract

Diversity is a long-studied topic in information retrieval that usually refers to the requirement that retrieved results should be non-repetitive and cover different aspects. In a conversational setting, an additional dimension of diversity matters: an engaging response generation system should be able to output responses that are diverse and interesting. Sequence-to-sequence (Seq2Seq) models have been shown to be very effective for response generation. However, dialogue responses generated by Seq2Seq models tend to have low diversity. In this paper, we review known sources and existing approaches to this low-diversity problem. We also identify a source of low diversity that has been little studied so far, namely model over-confidence. We sketch several directions for tackling model over-confidence and, hence, the low-diversity problem, including confidence penalties and label smoothing.

References Powered by Scopus

View more at Scopus

Cited by Powered by Scopus

View more at Scopus

Cite

CITATION STYLE

APA

Jiang, S., & de Rijke, M. (2018). Why are Sequence-to-Sequence Models So Dull? Understanding the Low-Diversity Problem of Chatbots. In Proceedings of the 2018 EMNLP Workshop SCAI 2018: The 2nd International Workshop on Search-Oriented Conversational AI (pp. 81–86). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/w18-5712

Readers' Seniority

PhD / Post grad / Masters / Doc 38

70%

Researcher 10

19%

Professor / Associate Prof. 3

Lecturer / Post doc 3

Readers' Discipline

Computer Science 49

78%

Engineering 6

10%

Linguistics 6

10%

Medicine and Dentistry 2

Why are Sequence-to-Sequence Models So Dull? Understanding the Low-Diversity Problem of Chatbots

Abstract

References Powered by Scopus

Rethinking the Inception Architecture for Computer Vision

Show and tell: A neural image caption generator

Learning discourse-level diversity for neural dialog models using conditional variational autoencoders

Cited by Powered by Scopus

Molecule Edit Graph Attention Network: Modeling Chemical Reactions as Sequences of Graph Edits

Non-parametric adaptation for neural machine translation

Improving neural response diversity with frequency-aware cross-entropy loss

Register to see more suggestions

Cite

Readers' Seniority

Readers' Discipline