arXiv:2510.08138cs.CVcs.AI2025-10中稿 · CVPR

提升视频语言模型的时序逻辑一致性,解决问答自相矛盾问题。

Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminability

  • 通过增强跨模态注意力对不同时刻视频标记的区分能力,改进时序理解。
  • 在多个数据集上使时序逻辑一致性提升15%-23%,且通用时序定位任务性能也改善。
  • 适用于需要准确时序推理的场景,如智能监控、视频摘要与交互式问答。

大型语言模型常生成自相矛盾的输出,严重影响其可靠性并阻碍实际应用。在视频语言模型(Video-LLMs)中,这一现象受到关注:模型对重述的问题无法提供逻辑一致的回答。然而其根本原因仍不清楚。本文采用可解释性驱动方法,统计分析并干预潜在因素。发现响应不一致的主要原因是跨模态注意力头难以有效区分不同时刻的视频标记。为此,提出一种注意力增强方法——时序条件注意力锐化(TCAS),基于注意力差异构建优化目标,以提升模型的时序分辨能力,从而增强时序逻辑一致性。实验表明,该方法显著提升视频语言模型的时序逻辑一致性。进一步分析证实,该方法确实增强了注意力头的时序可区分性。此外,该方法在通用视频时序定位任务中也取得性能提升,表明时序逻辑一致性是时序理解的关键因素。

原文摘要 · Abstract (English)

Large language models (LLMs) often generate self-contradictory outputs, which severely impacts their reliability and hinders their adoption in practical applications. In video-language models (Video-LLMs), this phenomenon recently draws the attention of researchers. Specifically, these models fail to provide logically consistent responses to rephrased questions based on their grounding outputs. However, the underlying causes of this phenomenon remain underexplored. In this work, we adopt an interpretability-driven approach to analyze, statistically summarize, and intervention the potential factors of the phenomenon. We find that one of the primary reasons for the inconsistency in responses lies in the inability of cross-modal attention heads to effectively distinguish video tokens across different timestamps. To address this, we propose an attention enhancement method called Temporally Conditioned Attention Sharpening (TCAS), which constructs an enhancement objective based on attention distinctions to enhance the model's temporal resolution capability, thereby improving its temporal understanding logic consistency. Experimental results demonstrate that our method significantly enhances the temporal logic consistency of Video-LLMs. Further analyses reveal that our method indeed improves the temporal discriminability of attention heads, validating our conclusions. Additionally, our method even achieves performance improvements in general video temporal grounding tasks, suggesting that temporal logic consistency is an important factor in temporal understanding.

视频语言模型时序一致性注意力机制可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。