让视频大模型能用文字和图像双向互动,实现精准对象指代与定位。
Object-centric Video Question Answering with Visual Grounding and Referring
- 引入视觉引导与指代机制,支持文本与图像双模态交互
- 提出时空叠加模块,将任意帧的视觉提示传播至全视频
- 构建新数据集VideoInfer,专攻对象中心的多轮视频问答与分割
视频大语言模型(VideoLLMs)在通用视频理解上已取得显著进展,但现有模型主要关注高层语义理解,仅支持纯文本输出,难以实现以对象为中心的多轮交互。本文提出三项贡献:(i) 构建支持输入时对象指代、输出时视觉定位的视频大模型,使用户可通过文本或图像双重方式与视频交互;(ii) 提出时空叠加模块(STOM),可将任意单帧的视觉提示传播至整个视频其余帧;(iii) 构建人工标注的视频指令数据集VideoInfer,包含需推理的对象中心问答对。在12个基准上的6项任务实验表明,所提模型在视频问答与目标分割任务中均持续优于基线,验证了其在多模态、对象中心视频与图像理解中的鲁棒性。
原文摘要 · Abstract (English)
Video Large Language Models (VideoLLMs) have recently demonstrated remarkable progress in general video understanding. However, existing models primarily focus on high-level comprehension and are limited to text-only responses, restricting the flexibility for object-centric, multiround interactions. In this paper, we make three contributions: (i) we address these limitations by introducing a VideoLLM model, capable of performing both object referring for input and grounding for output in video reasoning tasks, i.e., allowing users to interact with videos using both textual and visual prompts; (ii) we propose STOM (Spatial-Temporal Overlay Module), a novel approach that propagates arbitrary visual prompts input at any single timestamp to the remaining frames within a video; (iii) we present VideoInfer, a manually curated object-centric video instruction dataset featuring questionanswering pairs that require reasoning. We conduct comprehensive experiments on VideoInfer and other existing benchmarks across video question answering and referring object segmentation. The results on 12 benchmarks of 6 tasks show that our proposed model consistently outperforms baselines in both video question answering and segmentation, underscoring its robustness in multimodal, object-centric video and image understanding. Project page: https://qirui-chen.github.io/RGA3-release/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。