为语音大模型解释难题提供音素对齐新方法,显著降低计算负担。
SGPA: Spectrogram-Guided Phonetic Alignment for Feasible Shapley Value Explanations in Multimodal Large Language Models
- 通过频谱引导的音素对齐,将音频帧聚合成语义稳定的词段。
- 在语音基准测试中减少43倍模型评估次数,提升可解释性可行性。
- 适合需要高效音频解释的语音模型研究者与开发者使用。
通过Shapley值归因解释端到端语音语言模型的行为,在原始分词下难以实现:典型语句产生超过150个编码器帧,使联盟空间扩大约10^42倍;单个音频帧无独立语义;分割音素过渡的词边界引入掩码伪影。我们提出频谱引导的音素对齐(SGPA),一种四阶段流程,结合连接时序分类强制对齐与谱边界优化,生成声学稳定、词对齐的音频片段。在LFM2-Audio-1.5B和VoiceBench上的受控诊断显示,SGPA使模型评估次数减少43倍。统计检验表明,SGPA显著改变归因集中度,同时保持全局累积分布,确立其作为语音可解释性可行性的关键支撑层。
原文摘要 · Abstract (English)
Explaining the behavior of end-to-end audio language models via Shapley value attribution is intractable under native tokenization: a typical utterance yields over $150$ encoder frames, inflating the coalition space by roughly $10^{42}$ relative to text; individual audio frames lack standalone meaning; and token boundaries that bisect phonetic transitions introduce masking artifacts. We introduce Spectrogram-Guided Phonetic Alignment (SGPA), a four-stage pipeline that combines Connectionist Temporal Classification forced alignment with spectral boundary refinement to produce acoustically stable, word-aligned audio segments. Controlled diagnostics on LFM2-Audio-1.5B with VoiceBench show that SGPA yields a 43$\times$ reduction in model evaluations. Statistical testing confirms that SGPA significantly alters attribution concentration while preserving the global cumulative profile, establishing it as a feasibility-enabling layer for audio explainability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。