提出音频对话理解新评测框架,揭示文本误导问题并用音频孪生提升模型推理可靠性。
When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue

- 通过冲突问答对识别文本与语音信号的不一致,构建可控评测基准。
- 强文本模型在一致场景准确率超90%,但在冲突场景骤降至33-48%。
- 引入音频孪生表示,让模型显式利用语音线索,减少对文本的依赖陷阱。
理解口语对话需联合推理词汇内容与语调、说话风格等副语言信号。现有评估常允许仅基于文本转录或单模态解法的捷径,掩盖了模型是否真正基于语音进行判断。本文将此失败模式形式化为跨模态分歧:转录文提示看似合理但错误的表层解释,而声学线索如语调支持不同答案。我们提出可扩展框架,识别文本偏差的表层解释,并将分歧区域转化为冲突问答样本。同时包含转录与语音一致的情况,实现超越对抗性音频依赖的评估。由此构建了包含501个问题的ContraTalk基准,覆盖交互行为、情绪状态、对话行为、社会立场和对话意图五个维度。进一步开发了代理式推理框架,将语音转换为可读文本的音频孪生,暴露局部声学证据给推理模型。实验显示,强文本大模型在一致情况下准确率超过90%,但在冲突情况下下降至33-48%;直接音频大模型仅部分实现语音对齐,仍约30-40%选择文本偏差陷阱。音频孪生框架提升冲突情况准确率并降低陷阱选择,但一致情况表现仍依赖基础模型。结果揭示文本捷径是语音对话理解的重要失效模式,表明显式声学证据聚合可提供更可控的诊断与改进接口。
原文摘要 · Abstract (English)
Understanding spoken dialogue requires joint reasoning over lexical content and paralinguistic acoustic signals such as emotion and conversational intent. However, existing evaluations often allow shortcuts based on transcripts or single-modality solutions, obscuring whether models genuinely ground predictions in speech. We formalize this failure mode as cross-modal disagreement, where transcripts suggest plausible but incorrect surface interpretations while acoustic cues such as prosody or speaking style support different answers. We develop a scalable framework that identifies text-biased surface interpretations and converts disagreement regions into conflict QA examples. We also include consistent cases where transcript-based and speech-grounded interpretations agree, enabling evaluation beyond adversarial audio dependence. This results in ContraTalk, a controlled benchmark containing 501 questions across five discourse dimensions: interaction behavior, emotion state, dialogue act, social stance, and conversational intent. We further develop an agentic-style reasoning framework that converts speech into an Audio Twin, a text-readable representation of localized acoustic cues that exposes acoustic evidence to the reasoning model. Experiments show that strong text-only LLMs exceed 90% accuracy in consistent cases but drop to 33-48% in conflict cases. Direct AudioLLMs provide only partial grounding, still selecting the transcript-biased trap in roughly 30-40% of conflict cases. Our Audio Twin framework improves conflict-case accuracy while reducing trap selection, but its consistent-case behavior remains backbone-dependent. These results identify transcript-based shortcuts as an important failure mode in spoken dialogue understanding and show that explicit acoustic evidence aggregation provides a more controllable interface for diagnosing and improving speech-grounded reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。