arXiv:2510.12116cs.CLcs.AI2025-10EMNLP被引 12

揭示语音与文本输入在大模型中的差异机制,提出改进语音理解的新方法。

Understanding the Modality Gap: An Empirical Study on the Speech-Text Alignment Mechanism of Large Speech Language Models

  • 通过分析深层表示方向与幅度变化,发现语音与文本表征存在方向一致但幅度差异的矛盾。
  • 提出对齐路径得分量化词级对齐质量,其与模态差距强相关。
  • 针对关键词采用角度投影与长度归一化,显著提升语音输入正确性。

端到端大语音语言模型(LSLM)虽具备出色对话生成能力,但在语义理解基准测试中仍落后于传统流水线系统。本文通过系统实验发现,尽管语音-文本对齐训练后文本输入性能有所下降,但语音与文本输入间的性能差距更为显著,称为模态差距。我们分析了粗粒度与细粒度的文本与语音表征:在粗粒度层面,深层表示的方向(余弦相似度)逐渐对齐,但幅度(欧氏距离)持续发散;表征相似性与模态差距强相关。在细粒度层面,观察到自发的词级对齐模式,据此引入对齐路径得分以量化词级对齐质量,其与模态差距相关性更强。基于此,我们设计针对关键词的角度投影与长度归一化干预策略,有效提升了语音输入的正确性。本研究首次系统地揭示了LSLM中模态差距与对齐机制,为后续优化提供理论与方法指导。

原文摘要 · Abstract (English)

End-to-end Large Speech Language Models (LSLMs) have demonstrated impressive conversational generation abilities, yet consistently fall short of traditional pipeline systems on semantic understanding benchmarks. In this work, we reveal through systematic experimentation that although LSLMs lose some text input performance after speech-text alignment training, the performance gap between speech and text inputs is more pronounced, which we refer to as the modality gap. To understand this gap, we analyze both coarse- and fine-grained text and speech representations. At the coarse-grained level, representations of speech and text in deeper layers are found to be increasingly aligned in direction (cosine similarity), while concurrently diverging in magnitude (Euclidean distance). We further find that representation similarity is strongly correlated with the modality gap. At the fine-grained level, a spontaneous token-level alignment pattern between text and speech representations is observed. Based on this, we introduce the Alignment Path Score to quantify token-level alignment quality, which exhibits stronger correlation with the modality gap. Building on these insights, we design targeted interventions on critical tokens through angle projection and length normalization. These strategies demonstrate the potential to improve correctness for speech inputs. Our study provides the first systematic empirical analysis of the modality gap and alignment mechanisms in LSLMs, offering both theoretical and methodological guidance for future optimization.

语音理解模态对齐大模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。