通过分步分析揭示多语言语音文本模型的跨模态计算机制差异。
Generation-Step-Aware Framework for Cross-Modal Representation and Control in Multilingual Speech-Text Models
- 按生成步骤识别模态特异神经元,分解其表征与控制功能。
- 语音到文本的表征共享仅5-7%,而控制信号随生成阶段增强。
- 适用于研究跨模态对齐机制,尤其适合多语言模型开发者。
多语言语音-文本模型依赖跨模态语言对齐实现知识迁移,但其计算是否共享或模态特异尚不明确。本文提出一种生成步骤感知的评估框架,能(i)在不同解码步骤中识别各模态的语言选择性神经元,(ii)将其分解为语言表征与语言控制功能,(iii)通过重叠度量和因果干预实现跨模态比较。应用于SeamlessM4T,并在仅解码器模型Qwen2-Audio上验证通用性,发现跨模态表征对齐高度依赖模型:SeamlessM4T在生成过程中表现出显著的步骤依赖性,仅5-7%的语言表征神经元跨模态共享,且重叠从同语种向类型学相关语言转移;而Qwen2-Audio则保持较高跨模态共享(45-47%),并维持稳定强语言特异性对齐。相反,后期解码步骤中,SeamlessM4T的语言控制神经元展现出更强的语音到文本跨模态迁移。结果表明,生成步骤感知分析可揭示跨模态计算中既共享又模型特有的模式。
原文摘要 · Abstract (English)
Multilingual speech-text models rely on cross-modal language alignment to transfer knowledge between speech and text, but it remains unclear whether this reflects shared computation for the same language or modality-specific processing. We introduce a generation-step-aware framework for evaluating cross-modal computation that (i) identifies language-selective neurons for each modality at different decoding steps, (ii) decomposes them into language-representation and language-control roles, and (iii) enables cross-modal comparison via overlap measures and causal intervention. Applying our framework to SeamlessM4T, and additionally evaluating its generality on the decoder-only model Qwen2-Audio, we find that cross-modal language-representation alignment is highly model dependent. SeamlessM4T exhibits pronounced generation-step-dependent specialization, where only 5-7% of language-representation neurons are shared across modalities and overlap shifts from same-language to typologically related languages during autoregressive generation, whereas Qwen2-Audio maintains substantially larger cross-modal sharing (45-47%) together with stable, strongly language-specific alignment across decoding steps. In contrast, language-control neurons identified at later decoding steps exhibit progressively stronger cross-modal transfer from speech to text in SeamlessM4T. Together, these results demonstrate that generation-step-aware analysis can reveal both shared and model-specific patterns of cross-modal computation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。