arXiv:2606.06559cs.SDcs.AI2026-06

提出轻量级模块IRAFT,让语音对话系统在嘈杂环境中更抗干扰。

IRAF: Interference-Resilient Adaptive Fusion for Noise-Robust End-to-End Full-Duplex Spoken Dialogue Systems

论文配图:IRAF: Interference-Resilient Adaptive Fusion for Noise-Robust End-to-End Full-Duplex Spoken Dialogue Systems
图 1 · 摘自论文原文
  • 通过实时评估用户语音可信度,动态调节其对大模型的贡献
  • 在双说话人干扰下,响应质量与交互稳定性显著提升
  • 适合部署于实时语音助手等需要全双工交互的场景

全双工语音对话模型允许语音助手同时听和说,实现自然的实时重叠交互。然而,在真实声学环境下,端到端双通道模型可能因干扰说话人声音泄漏到用户麦克风中而性能下降:这些干扰信号会被编码为用户输入的一部分,污染大语言模型的条件输入,导致话轮切换不稳定和响应质量降低。本文提出抗干扰自适应融合(IRAF),一种轻量级、支持流式处理的模块,可逐帧调节用户音频对大模型的贡献。IRAFT从目标说话人和用户音频嵌入中预测一个标量可靠性门控值,并在融合前重新缩放用户表示。在MS-MARCO和InstructS2S-200K数据集上的实验表明,该方法在存在干扰说话人条件下持续提升了响应质量和全双工交互表现。

原文摘要 · Abstract (English)

Full-duplex spoken dialogue models allow voice agents to listen and speak concurrently, enabling natural interaction with real-time overlap. However, end-to-end dual-channel models that jointly encode user and agent streams may degrade in realistic acoustic environments: interfering speakers leaking into the user microphone can be encoded as part of the user query, corrupting the LLM's conditioning and causing unstable turn-taking and reduced response quality. We propose Interference-Resilient Adaptive Fusion (IRAF), a lightweight, streaming-compatible module that modulates the contribution of user audio to the LLM frame by frame. IRAF predicts a scalar reliability gate from target-speaker and user audio embeddings and rescales user representations before fusion with agent embeddings. Experiments on MS-MARCO and InstructS2S-200K show consistent gains in response quality and full-duplex interaction under interfering-speaker conditions.

语音对话抗干扰全双工大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。