arXiv:2607.21943cs.SD2026-07

让语音模型真正听懂乱场中的对话,而非靠提示抄答案。

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding

论文配图:Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding
图 1 · 摘自论文原文
  • 用音频线索引导模型听清说话人,不给答案文本
  • 多场景测试下误识别率从25%-71%降至9%-15%
  • 适合需要真实语音理解的智能助手、会议记录等场景

通用语音模型在清晰单人说话时表现良好,但在多人重叠、嘈杂环境下的准确率急剧下降,而这正是需要区分谁说了什么的场景。直接添加简短场景描述看似有效,实则导致模型抄袭提示文本而非真实聆听,造成虚假高分。我们称此为‘感知绕过’问题,并提出音频锚定的支架上下文(AGSC)来解决。AGSC分三步:首先从音频中构建引导聆听的线索而不暴露答案;其次通过答案重叠和静音测试检测线索泄露与音频依赖性;最后将这些线索用于训练但测试时消失,实现无提示能力。在三个异构的Omni模型上,使用AGSC训练后,在重叠嘈杂语音上的无提示上限平均词错误率(mpWER)从25%-71%降至9%-15%。针对流式处理,我们设计联合GDPO任务,使模型学会何时使用线索,并从独立归一化的门控、格式和转录奖励中生成带说话人标注的转录。内部化后,推理开销几乎不变。

原文摘要 · Abstract (English)

Omni models transcribe clean, single-speaker speech well, but their accuracy drops sharply when speakers overlap and the scene is noisy, exactly where knowing who said what matters most. A natural fix is a short scene description. We show why this is risky: answer-bearing text lets the model copy instead of listen, so the score rises although nothing has been heard; a silent test exposes this shortcut at once. We call this failure mode perception bypass and address it with Audio-Grounded Scaffold Context (AGSC). AGSC links three steps: first, we build clues from audio to guide listening without giving the answer; second, answer-overlap and silence tests probe them for leakage and audio dependence; finally, those clues scaffold training but vanish at test time, yielding no-clue capability. Across three heterogeneous Omni models, training on AGSC lowers no-clue capped mean permutation word error rate (mpWER) on overlapping, noisy speech from 25%-71% to 9%-15%. For streaming control, we formulate a joint GDPO task in which the model learns when to use a clue and how to produce a speaker-attributed transcript from separately normalized format, gate, and transcript rewards. After internalization, AGSC adds almost no inference overhead.

语音理解多说话人音频引导模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。