ConvMix: A Mixed-Criteria Data Augmentation Framework for Conversational Dense Retrieval
Published in Proceedings of the 40th AAAI Conference on Artificial Intelligence (AAAI 2026), 2026
Conversational dense retrieval depends on relevance judgments that reveal a userβs search intent across multiple context-dependent turns, but suitable training data are scarce and costly to annotate. ConvMix introduces a mixed-criteria augmentation framework that uses large language models to generate scalable, two-sided relevance judgments from multiple perspectives. Semantic diversity controls, noise filtering, and near-distribution supervision help maintain sample quality and reduce distribution mismatch when human-annotated and generated data are combined. Evaluations on five conversational search benchmarks demonstrate consistent improvements over existing augmentation approaches.
π In Proceedings of the 40th AAAI Conference on Artificial Intelligence (AAAI 2026), Volume 40, Issue 18, pp. 15555β15563.
π Paper Link (AAAI)
