arXiv:2607.25355cs.SD2026-07

揭秘音频模型微调如何让语言解码器更精准读取时间线索。

From Semantics to Readout: Mechanistic Understanding of Audio Tokens after Fine-Tuning for Temporal Audio Grounding

论文配图:From Semantics to Readout: Mechanistic Understanding of Audio Tokens after Fine-Tuning for Temporal Audio Grounding
图 1 · 摘自论文原文
  • 通过四类分析揭示音频令牌在微调后的语义与可读性变化。
  • 微调后解码器对事件信息的访问能力显著提升,尤其在早期和中间层。
  • 适合关注多模态模型可解释性与时间定位任务的研究者。

大型音频-语言模型(LALMs)通过原生音频令牌向语言解码器传递声学证据,但这些令牌的内部作用机制尚不清晰。以时间音频定位为诊断场景,本文通过四种互补分析——查询条件下的令牌语义、校准的令牌读出、时间窗探测及生成过程中的残差增量擦除——研究了语言模型微调对原生音频令牌状态的分层语义、解码器可访问性及时间输出对齐的影响。结果显示,在微调后,时间定位性能显著提升;对Qwen2.5-Omni的语义分析表明,目标事件的潜在证据在微调前已存在,且与查询事件最相关的时间位置在微调前后保持一致。微调后,音频令牌中事件相关信息对解码器的可访问性增强,尤其在早期和中间层,跨检查点对照显示这一改进主要源于解码器适应。时间探测表明基础模型已包含可恢复的标注窗口信息,微调主要提升与各检查点自身预测时间支持的对齐度。残差增量擦除进一步显示,移除预测窗口内的音频令牌更新比随机移除相同数量的更新更严重损害时间戳生成。Qwen2-Audio也观察到类似的解码器可读性与预测对齐性提升。综合结果支持‘从语义到读出’的解释框架:微调帮助解码器更可靠地读取已有事件证据并连接至时间输出。

原文摘要 · Abstract (English)

Large audio-language models (LALMs) convey acoustic evidence to language decoders through native audio tokens, yet the internal roles of these tokens remain poorly understood. Using temporal audio grounding as a diagnostic setting, we examine how language-model fine-tuning affects the layerwise semantics, decoder accessibility, and temporal output alignment of native audio-token states through four complementary analyses: query-conditioned token semantics, calibrated token readout, temporal-window probes, and residual-delta erasure during generation. Alongside substantial improvements in temporal localization, semantic analysis of Qwen2.5-Omni shows that latent evidence for queried events is already present before fine-tuning and that the audio tokens most strongly aligned with the queried event appear at similar temporal positions before and after fine-tuning. After fine-tuning, event-related information in audio tokens becomes more accessible to the decoder, especially in early and middle layers, and a cross-checkpoint control shows that this improvement arises primarily from decoder adaptation. Temporal probes show that the base checkpoint already contains recoverable information about annotated windows and that fine-tuning mainly improves alignment with each checkpoint's own predicted temporal support. Residual-delta erasure further shows that removing audio-token updates within predicted windows harms timestamp generation more than removing the same number of randomly selected updates. The same broad improvements in decoder readability and prediction alignment also appear in Qwen2-Audio. Together, these results support a semantics-to-readout account in which grounding fine-tuning helps the decoder read existing event evidence and connect it more reliably to temporal outputs.

音频理解多模态可解释性时间定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。