构建首个中文语音消歧数据集,揭示语音如何澄清文本歧义。
DEBATE: A Dataset for Disentangling Textual Ambiguity in Mandarin Through Speech
- 采集1001个含歧义的中文句子,由10位母语者录音,分析语音线索
- 发现机器对语音意图的理解远落后于人类,性能差距显著
- 适合研究语音消歧、人机交互与多模态语言理解的学者使用
尽管文本和视觉消歧已有大量研究,语音消歧(DTS)仍被严重忽视,主要因缺乏高质量的语音-文本配对数据集。为此,我们提出DEBATE,一个公开的中文语音-文本数据集,用于研究语音线索(发音、停顿、重音、语调)如何化解文本歧义并揭示说话人真实意图。该数据集包含1,001个精心挑选的歧义语句,每条由10名母语者录制,涵盖多样化的语言歧义及其语音消解方式。我们详细描述了数据收集流程,并进行严格的质量分析。此外,我们对三种先进的大模型进行了基准测试,结果表明机器在理解语音意图方面与人类存在明显且巨大的性能差距。DEBATE是该领域的首次尝试,为跨语言、跨文化的语音消歧数据集构建提供了基础。数据集及代码已开源:https://github.com/SmileHnu/DEBATE。
原文摘要 · Abstract (English)
Despite extensive research on textual and visual disambiguation, disambiguation through speech (DTS) remains underexplored. This is largely due to the lack of high-quality datasets that pair spoken sentences with richly ambiguous text. To address this gap, we present DEBATE, a unique public Chinese speech-text dataset designed to study how speech cues and patterns-pronunciation, pause, stress and intonation-can help resolve textual ambiguity and reveal a speaker's true intent. DEBATE contains 1,001 carefully selected ambiguous utterances, each recorded by 10 native speakers, capturing diverse linguistic ambiguities and their disambiguation through speech. We detail the data collection pipeline and provide rigorous quality analysis. Additionally, we benchmark three state-of-the-art large speech and audio-language models, illustrating clear and huge performance gaps between machine and human understanding of spoken intent. DEBATE represents the first effort of its kind and offers a foundation for building similar DTS datasets across languages and cultures. The dataset and associated code are available at: https://github.com/SmileHnu/DEBATE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。