arXiv:2507.18531cs.CV2025-07被引 13

让视频描述更精准:通过双策略填补视觉模型时空理解鸿沟

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning

  • 用提示组合建模用户意图与视频的隐含关联
  • 引入轻量级框适配器增强对象语义先验信息
  • 显著提升多模型对目标细节的精准控制能力

意图导向的可控视频描述旨在根据用户自定义意图,为视频中的特定目标生成目标化描述。当前大型视觉语言模型(LVLMs)虽具备强大的指令跟随与视觉理解能力,但在时间序列中实现细粒度的空间控制仍存在明显短板,导致难以精确响应用户意图。为此,我们提出IntentVCNet,从提示和模型两方面统一LVLM固有的时空理解知识,以弥合这一关键的时空鸿沟。具体而言,我们设计了一种提示组合策略,使大语言模型能够建模用户意图提示与视频序列之间的隐含关系;同时提出一种参数高效的目标框适配器,将对象语义信息注入全局视觉上下文,使视觉标记具备用户意图先验。实验表明,两种策略结合能显著增强模型在视频序列中建模空间细节的能力,有效支持准确生成意图导向的可控描述。该方法在多个开源LVLM上达到最优性能,并在IntentVC挑战赛中位列第二。代码已开源。

原文摘要 · Abstract (English)

Intent-oriented controlled video captioning aims to generate targeted descriptions for specific targets in a video based on customized user intent. Current Large Visual Language Models (LVLMs) have gained strong instruction following and visual comprehension capabilities. Although the LVLMs demonstrated proficiency in spatial and temporal understanding respectively, it was not able to perform fine-grained spatial control in time sequences in direct response to instructions. This substantial spatio-temporal gap complicates efforts to achieve fine-grained intention-oriented control in video. Towards this end, we propose a novel IntentVCNet that unifies the temporal and spatial understanding knowledge inherent in LVLMs to bridge the spatio-temporal gap from both prompting and model perspectives. Specifically, we first propose a prompt combination strategy designed to enable LLM to model the implicit relationship between prompts that characterize user intent and video sequences. We then propose a parameter efficient box adapter that augments the object semantic information in the global visual context so that the visual token has a priori information about the user intent. The final experiment proves that the combination of the two strategies can further enhance the LVLM's ability to model spatial details in video sequences, and facilitate the LVLMs to accurately generate controlled intent-oriented captions. Our proposed method achieved state-of-the-art results in several open source LVLMs and was the runner-up in the IntentVC challenge. Our code is available on https://github.com/thqiu0419/IntentVCNet.

视频生成意图控制视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。