arXiv:2506.21116cs.CVcs.AI2025-06

解决视频模型在多镜头场景下的身份混淆问题

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes

  • 用实例提示注入机制增强跨场景身份识别
  • 在新数据集MultiClip-Bench上准确率提升12.7%
  • 适合需要精准跟踪多镜头视频的场景研究者

视频大语言模型虽具强大理解能力,但在多镜头场景(如视角变化或场景切换)下易出现实例身份丢失和关键帧遗漏。本文指出现有数据集缺乏多镜头标注,因此构建了新数据集MultiClip-Bench,包含密集描述与基于指令的问答对。实验发现,训练集显著提升多镜头性能,测试基准可靠评估模型能力。进一步分析表明,当前模型仅以离散或有损方式编码实例特征,易丢失身份信息。为此提出IPFormer-VideoLLM,通过高效的注意力连接器将实例级特征作为实例提示注入,实现跨场景实例信息聚合。实验证明,所提数据集与模型显著提升多场景视频理解能力,并在多个视频基准上展现优势。

原文摘要 · Abstract (English)

Video Large Language Models (VideoLLMs) have demonstrated remarkable understanding capabilities, but are found struggling to tackle multi-shot scenarios,e.g., video clips with varying camera angles or scene changes. This challenge can render failures such as instance identity forgetting and key frame negligence. In this work, we first attribute the challenge to the lack of multi-shot annotations among existing datasets and therefore we introduce a new dataset termed MultiClip-Bench, featuring dense descriptions and instruction-based question-answering pairs tailored for multi-shot scenarios. We empirically find that the training set significantly boosts the multi-shot performance, while the testing benchmark provides a reliable measure of the model capability in multi-shot scenarios. By further analyzing and discovering that current models only encode instance features in a discrete or lossy manner, at the risk of missing identity information, we then contribute a new model IPFormer-VideoLLM. Its key idea is the injection of instance-level features as instance prompts through an efficient attention-based connector. This allows for the aggregation of instance-specific information across scenes. Experiments demonstrate that our proposed dataset and model not only enhance the multi-scene video understanding significantly, but also offer distinct advantages across various video benchmarks.

视频理解多镜头提示注入视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。