arXiv:2602.00683cs.CV2026-02被引 1

显式建模视频时间关系,提升模型对动态内容的理解能力

Video Understanding: Through A Temporal Lens

论文配图:Video Understanding: Through A Temporal Lens
图 1 · 摘自论文原文
  • 用大视觉语言模型与抗噪对比学习自动标注视频
  • 提出循环适配器,在少量数据下有效捕捉时间动态
  • 引入状态空间层并构建新基准,支持长视频高效建模

本论文围绕如何利用视频元素间的时间关系推进视频理解这一核心问题展开研究。针对现有方法的局限性,提出五项贡献:(1) 基于大视觉语言模型与抗噪对比学习(带减法角度边距)的自动标注框架;(2) 使用‘循环适配器’的参数高效微调策略,以在低数据场景下捕捉时间动态;(3) 集成状态空间层(SSL)实现长视频高效建模,并引入两个新长期基准——用于第一人称视角和长片内容;(4) 设计新型对比学习框架,显式建模运动与视频片段间的细粒度关联;(5) 对大视觉语言模型(LVLMs)进行综合实证研究,发现视觉-语言接口是时间推理瓶颈,据此提出面向扩展视频理解的‘时间导向配方’。上述工作共同表明,显式时间建模显著增强模型对视频流体特性的表征与推理能力。

原文摘要 · Abstract (English)

This thesis explores the central question of how to leverage temporal relations among video elements to advance video understanding. Addressing the limitations of existing methods, the work presents a five-fold contribution: (1) an automatic annotation framework that utilizes large vision-language models and a noise-robust contrastive learning objective with a subtractive angular margin; (2) a parameter-efficient fine-tuning strategy using "recurrent adapters" to capture temporal dynamics in low-data regimes; (3) the integration of State Space Layers (SSL) for efficient long-form video modeling, supported by the introduction of two new long-term benchmarks for egocentric and feature-length content; (4) a novel contrastive learning framework designed to explicitly model fine-grained relations between motions and video moments; and (5) a comprehensive empirical study on Large Vision-Language Models (LVLMs) that identifies the visual-language interface as a bottleneck for temporal reasoning, leading to a new "temporal-oriented recipe" for upscaled video understanding. Collectively, these contributions demonstrate that explicit temporal modeling significantly enhances a model's ability to represent and reason about the fluid nature of video content.

视频理解时间建模大模型对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。