揭示语音模型语义冗余导致的模态差距,提出深层表示演化新视角。
Anatomy of the Modality Gap: Dissecting the Internal States of End-to-End Speech LLMs
- 通过跨层CKA分析语音与文本表征演化路径
- 发现语音表征存在宽幅跨层对齐带,源于语义跨帧冗余
- 指出单纯统计校准无效,需在词元/时序粒度改进
近期大语音语言模型显著缩小了声学信号与语言理解之间的差距,但在语音输入任务中的表现仍落后于直接文本推理。本文超越静态几何对齐,分析语音与文本表征在多层间的动态演化过程。我们在SpeechMMLU和VoiceBench BBH上评估了四种开源端到端模型,采用带语音-文本词元对齐的跨层CKA分析,发现语音表征呈现宽幅跨层对齐带,归因于语音中语义内容跨多个帧分布的冗余性。该对齐模式在不同分析配置下结构稳定。关键的是,仅在输入层进行简单统计校准不仅无效,甚至有害,表明模态差距并非单纯的分布偏移。整体结果提示瓶颈在于将冗余语音压缩为稳定的晚期决策,推动未来工作应在词元或时间粒度而非特征层面进行优化。
原文摘要 · Abstract (English)
Recent advancements in Large Speech-Language Models have significantly bridged the gap between acoustic signals and linguistic understanding. However, a persistent performance disparity remains in speech-based input tasks compared to direct text inference. In this paper, we investigate the dynamic roots of this modality gap beyond static geometric alignment, analyzing how speech and text representations evolve layer-by-layer. We evaluate four open-weight end-to-end models on SpeechMMLU and VoiceBench BBH. Using cross-layer CKA analysis with speech-text token alignment, we find that speech representations exhibit a broad cross-layer alignment band, attributable to the redundant nature of speech where semantic content spans multiple frames. We show that these alignment patterns are structurally stable across different analysis configurations. Crucially, simple statistical calibration is insufficient and can be detrimental when applied at the input layer, indicating that the modality gap is not a mere distribution shift. Overall, our results suggest that the bottleneck lies in condensing redundant speech into stable late-layer decisions, motivating future solutions that operate at the token or temporal granularity instead of feature-level matching.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。