arXiv:2509.06335cs.CV2025-09中稿 · WACV 2026

用物体定位信息提升视频大模型的时间感知能力

Harnessing Object Grounding for Time-Sensitive Video Understanding

  • 引入物体定位信息,通过轻量模块实时编码物体特征
  • 在多个任务和数据集上超越原始模型与文本描述增强方案
  • 适合需要精准时间定位的视频理解研究者使用

我们提出利用物体定位(GO)来增强视频大语言模型(Video-LLM)的时间敏感视频理解(TSV)能力。实验表明,基于LITA模型的初步测试显示,帧内物体定位信息对时序定位任务有益。尽管在提示中加入物体文本描述可提升性能,但会增加令牌长度并受物体信息噪声影响。为此,我们设计了GO-Tokenizer,一个轻量级附加模块,利用现成物体检测器实时编码紧凑的物体信息。实验结果表明,预训练时使用GO-Tokenizer优于原生Video-LLM及其依赖物体文本描述的版本。该优势在不同模型、数据集和任务(如推理时序定位、密集描述)间均具泛化性。

原文摘要 · Abstract (English)

We propose to improve the time-sensitive video understanding (TSV) capability of video large language models (Video-LLMs) with grounded objects (GO). We hypothesize that TSV tasks can benefit from GO within frames, which is supported by our preliminary experiments on LITA, a state-of-the-art Video-LLM for reasoning temporal localization. While augmenting prompts with textual descriptions of these object annotations improves the performance of LITA, it also introduces extra token length and susceptibility to the noise in object-level information. To address this, we propose GO-Tokenizer, a lightweight add-on module for Video-LLMs leveraging off-the-shelf object detectors to encode compact object information on the fly. Experimental results demonstrate that pretraining with GO-Tokenizer outperforms the vanilla Video-LLM and its counterpart, utilizing textual descriptions of objects in the prompt. The gain generalizes across different models, datasets, and video understanding tasks, such as reasoning temporal localization and dense captioning.

视频理解物体定位时序推理大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。