解决语音模型中声学与语义干扰问题,提升全双工对话自然度。
Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs

- 分层参数分离设计,解耦声学与语义分支,减少梯度冲突。
- 在Spoken QA上提升7.4%,全双工流畅度提升28.5%。
- 首个揭示并解决全双工语音模型干扰根源的工作,适合语音交互研究者。
构建无缝、高性能的原生智能全双工语音语言模型(SLMs)仍是语音与自然语言处理领域的关键挑战。尽管取得进展,现有方法仍受严重模态干扰制约,导致知识退化与语义不连贯,使模型表现生硬且不智能。本文通过细致分析模型优化动态,揭示了根本原因:当声学与语义模态被迫共享深层参数空间时,二者间存在固有的梯度冲突。基于此洞察,提出Lychee-FD框架,采用分层参数分离策略,在深层解耦冲突模态,同时通过专用语义对齐通道保持跨模态一致性。在多个全双工基准测试上的实验证明,该方法显著超越现有技术,Spoken QA性能提升7.4%,全双工交互流畅度提升28.5%,且不牺牲推理效率。据我们所知,这是首个揭示并阐明全双工SLMs中模态干扰根源,并设计出优雅分层结构与实用解决方案的工作。
原文摘要 · Abstract (English)
Developing seamless, high-performance, native intelligent full-duplex Spoken Language Models (SLMs) remains a critical challenge and long-standing goal for the speech and NLP community. Despite notable progress, recent endeavors are fundamentally constrained by severe modality interference, which causes substantial knowledge degradation and compromises semantic integrity -- ultimately making full-duplex SLMs feel unnatural and unintelligent. In this paper, through an exhaustive fine-grained analysis of model optimization dynamics, we uncover the root cause of such performance degradation, revealing that modality interference arises from inherent gradient conflicts between acoustic and semantic modeling when the two modalities are forced to share a deep parameter space. Guided by this key insight, we introduce Lychee-FD, a native end-to-end full-duplex framework designed to mitigate modality interference. Importantly, we propose a hierarchical parameter separation strategy that decouples conflicting modalities in deep layers while preserving cross-modality coherence via a dedicated semantic alignment channel. Extensive experiments on multiple full-duplex benchmarks demonstrate that our method significantly advances the state of the art, yielding substantial improvements in both speech intelligence (+7.4% on Spoken QA) and full-duplex interaction fluidity (+28.5% on FullDuplexBench 1.5) without compromising inference efficiency. To the best of our knowledge, this work is the first to achieve two key advances: 1) uncovering and elucidating the root cause of modality interference in full-duplex SLMs, and 2) designing an elegant hierarchical model together with a practical solution for seamless, high-performance, native intelligent full-duplex SLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。