arXiv:2502.11594cs.CV2025-02ACL被引 14

提升视频模型对物体运动的精细感知能力,助力长时序理解。

iMOVE: Instance-Motion-Aware Video Understanding

  • 构建首个大规模实例级运动感知数据集iMOVE-IT,含丰富运动标注与时空互监督任务。
  • 提出iMOVE模型,在保持效率的同时精准捕捉实例时空运动细节,长时序理解性能显著提升。
  • 适合需要精细动作分析的场景,如行为识别、自动驾驶视频理解。

提升视频大语言模型对细粒度实例时空运动的感知能力,对增强其时间与通用视频理解至关重要。然而,现有模型难以捕捉详细且复杂的实例运动。为此,本文从数据与模型两方面进行改进:在数据层面,精心构建了首个大规模实例运动感知视频指令微调数据集iMOVE-IT,包含全面的实例运动标注及时空互监督任务,为模型提供充足训练支持;在此基础上,提出iMOVE——一种实例运动感知视频基础模型,采用事件感知时空高效建模机制,有效保留关键实例时空运动信息的同时保持计算高效,并引入相对时空位置标记以增强对实例时空位置的感知能力。评估结果表明,iMOVE不仅在视频时间理解与通用视频理解上表现优异,更在长时序视频理解方面展现出显著优势。

原文摘要 · Abstract (English)

Enhancing the fine-grained instance spatiotemporal motion perception capabilities of Video Large Language Models is crucial for improving their temporal and general video understanding. However, current models struggle to perceive detailed and complex instance motions. To address these challenges, we have made improvements from both data and model perspectives. In terms of data, we have meticulously curated iMOVE-IT, the first large-scale instance-motion-aware video instruction-tuning dataset. This dataset is enriched with comprehensive instance motion annotations and spatiotemporal mutual-supervision tasks, providing extensive training for the model's instance-motion-awareness. Building on this foundation, we introduce iMOVE, an instance-motion-aware video foundation model that utilizes Event-aware Spatiotemporal Efficient Modeling to retain informative instance spatiotemporal motion details while maintaining computational efficiency. It also incorporates Relative Spatiotemporal Position Tokens to ensure awareness of instance spatiotemporal positions. Evaluations indicate that iMOVE excels not only in video temporal understanding and general video understanding but also demonstrates significant advantages in long-term video understanding.

视频理解运动感知大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。