揭秘视频语言模型中时间线索的流动路径并提升其理解能力
CircuitProbe: Tracing Visual Temporal Evidence Flow in Video Language Models
- 通过分阶段分析框架追踪视觉时间证据在模型中的传播路径
- 发现特定注意力头对时间理解至关重要,干预后在时序任务上提升2.4%
- 适合关注视频理解机制与模型可解释性的研究者
自回归大视觉-语言模型(LVLM)通过将视频特征投影到语言模型嵌入空间生成连续视觉标记,实现视频与语言的交互。然而,时间证据在何处表示及其如何影响解码过程仍不清晰。为此,我们提出CircuitProbe——一个电路级分析框架,分为两个阶段:(i) 视觉审计,定位投影视频标记序列中的物体语义,并通过靶向消融和受控替换揭示其因果必要性;(ii) 语义追踪,利用logit-lens探测技术追踪层间物体与时间概念的演化,并结合时间帧干预评估对时间结构的敏感性。基于分析结果,我们设计了一种符合观察的外科式干预:识别时序特化的注意力头,并在语义追踪揭示的关键层区间内选择性增强其作用。该分析驱动的干预在时序密集型的TempCompass基准上实现了最高2.4%的绝对提升,验证了所提电路级分析在提升视频语言模型时间理解上的正确性、有效性与实用价值。
原文摘要 · Abstract (English)
Autoregressive large vision--language models (LVLMs) interface video and language by projecting video features into the LLM's embedding space as continuous visual token embeddings. However, it remains unclear where temporal evidence is represented and how it causally influences decoding. To address this gap, we present CircuitProbe, a circuit-level analysis framework that dissects the end-to-end video-language pathway through two stages: (i) Visual Auditing, which localizes object semantics within the projected video-token sequence and reveals their causal necessity via targeted ablations and controlled substitutions; and (ii) Semantic Tracing, which uses logit-lens probing to track the layer-wise emergence of object and temporal concepts, augmented with temporal frame interventions to assess sensitivity to temporal structure. Based on the resulting analysis, we design a targeted surgical intervention that strictly follows our observations: identifying temporally specialized attention heads and selectively amplifying them within the critical layer interval revealed by Semantic Tracing. This analysis-driven intervention yields consistent improvements (up to 2.4% absolute) on the temporal-heavy TempCompass benchmark, validating the correctness, effectiveness, and practical value of the proposed circuit-level analysis for temporal understanding in LVLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。