构建中英文语音对话模型评测基准,揭示复杂对话中的真实挑战
C3: A Bilingual Benchmark for Spoken Dialogue Models Exploring Challenges in Complex Conversations
- 设计包含1079个实例的双语语音对话数据集
- 通过大模型评估法实现接近人类判断的性能测试
- 聚焦歧义与上下文依赖等真实对话难题,适合对话系统研究者
语音对话模型(SDMs)近年来受到广泛关注,因其能直接对用户语音查询生成回应。然而,相较于受益于大量基准测试的文本大语言模型(LLMs),针对其在理解与模拟人类对话方面实际效果的研究仍显不足。语音交互固有复杂性远超文本,包括语义歧义(如多义词)、语音层面的异形同音、异音同形及重音差异,以及省略、指代和多轮交互等上下文依赖问题。为揭示当前SDM发展水平并应对这些挑战,本文提出一个包含1079个实例的中英文语音对话基准数据集,并配套基于大模型的评估方法,该方法能贴近人类判断,全面评估模型在处理复杂对话挑战时的表现。
原文摘要 · Abstract (English)
Spoken Dialogue Models (SDMs) have recently attracted significant attention for their ability to generate voice responses directly to users' spoken queries. Despite their increasing popularity, there exists a gap in research focused on comprehensively understanding their practical effectiveness in comprehending and emulating human conversations. This is especially true compared to text-based Large Language Models (LLMs), which benefit from extensive benchmarking. Human voice interactions are inherently more complex than text due to characteristics unique to spoken dialogue. Ambiguity poses one challenge, stemming from semantic factors like polysemy, as well as phonological aspects such as heterograph, heteronyms, and stress patterns. Additionally, context-dependency, like omission, coreference, and multi-turn interaction, adds further complexity to human conversational dynamics. To illuminate the current state of SDM development and to address these challenges, we present a benchmark dataset in this paper, which comprises 1,079 instances in English and Chinese. Accompanied by an LLM-based evaluation method that closely aligns with human judgment, this dataset facilitates a comprehensive exploration of the performance of SDMs in tackling these practical challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。