arXiv:2606.16472cs.CL2026-06

通过感知-执行对齐提升对话系统上下文忠实度

From Awareness to Adherence: Bridging the Context Gap in Spoken Dialogue Systems via Context-Aware Decoding

论文配图:From Awareness to Adherence: Bridging the Context Gap in Spoken Dialogue Systems via Context-Aware Decoding
图 1 · 摘自论文原文
  • 利用注意力机制识别关键历史回合,动态增强上下文信号
  • 在Audio MultiChallenge上,语义记忆与自一致性任务显著提升
  • 适合需要严格遵循对话历史的语音交互系统开发者

尽管端到端语音对话系统取得成功,但在多轮对话中保持严格的上下文忠实度仍是挑战。以往研究将失败归因于模型遗忘对话历史,但我们指出一个同样重要却被忽视的瓶颈:隐式上下文感知与主动遵循之间存在差距。尽管模型内部能识别相关历史话语,但强参数先验常在解码时压制这些信号。为此,我们提出音频适配的上下文感知解码(CAD)方法。通过内部分析注意力机制提取关键历史轮次,在推理时对比有无该关键上下文的输出分布,直接放大多模态上下文信号。在Audio MultiChallenge基准上的评估表明,该方法在语义记忆与自一致性子任务上均有显著提升,成功实现严格的上下文忠实响应。

原文摘要 · Abstract (English)

Despite the success of end-to-end (E2E) spoken dialogue systems, maintaining strict context adherence in multi-round conversations remains a challenge. While prior works attribute these failures to models forgetting dialogue history, we highlight an equally critical but overlooked bottleneck: a gap between latent context awareness and active adherence. Although models internally recognize relevant past utterances, strong parametric priors often overshadow these signals during decoding. To bridge this gap, we propose an audio-adapted Context-Aware Decoding (CAD) approach. By leveraging internal attention mechanisms to isolate key historical rounds, our approach contrasts output distributions with and without this key context during inference, directly amplifying multimodal contextual signals. Evaluations on the Audio MultiChallenge benchmark demonstrate significant improvements in Semantic Memory and Self Coherence subtasks, successfully enforcing strict, context-faithful adherence.

对话系统上下文感知语音交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。