用视觉信息修复嘈杂环境中的语音令牌,让对话系统更懂人话。
Noise Adaptive Streaming Audio-Visual Speech Token Enhancement for Robust Full-Duplex Spoken Dialogue Models

- 模块化前端实时融合音视频,修复噪声和重叠语音中的语义令牌。
- 在同数据集干扰下,对话连贯性评分从1.42提升至1.91。
- 无需训练主模型,适合快速部署到现有对话系统中。
全双工对话系统可同时听与说,但纯音频感知在背景噪声和重叠说话时表现不佳,导致回应不连贯。近年来的音视频对话方法表明,结合唇动等视觉线索可提升抗噪能力。然而,现有方法通常需对大型对话模型进行多模态训练,成本高昂。本文提出AV-STE,一种模块化流式音视频前端,能在语音大模型处理前,从噪声音频与唇部视频中恢复受损的语义语音令牌。下游对话模型保持完全冻结,保留其预训练对话能力。集成冻结版Moshi后,AV-STE在同数据集说话人干扰下,使GPT-4o评估的回应连贯性从1.42提升至1.91,且基本维持了自然对话轮次行为。该效果亦可迁移到跨领域无缝交互场景。
原文摘要 · Abstract (English)
Full-duplex spoken dialogue systems enable simultaneous listening and speaking, but their audio-only perception often fails under background noise and overlapping speech, leading to incoherent responses. Recent audio-visual dialogue approaches show that incorporating visual cues such as lip movements improve robustness under audio corruption. However, existing approaches often adapt the large speech dialogue model itself to process visual input, requiring costly multimodal training. We propose AV-STE, a modular streaming audio-visual front-end that restores corrupted semantic speech tokens from noisy audio and lip video before they reach the speech LLM. The downstream dialogue model remains entirely frozen, preserving its pretrained conversational capabilities. When integrated with frozen Moshi, AV-STE improves average GPT-4o-judged response coherence from 1.42 to 1.91 under same-dataset speaker interference while largely preserving turn-taking behavior. Gains also transfer to out-of-domain Seamless Interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。