构建首个大规模摄像机运动理解数据集,助力视频分析模型提升感知能力
Towards Understanding Camera Motions in Any Video
- 提出摄像机运动分类体系,区分语义与几何运动类型
- 发现现有模型在内容依赖型运动上表现差,几何轨迹估计不准确
- 提供标注指南与训练教程,适合视频理解与多模态研究者使用
我们提出CameraBench,一个包含约3,000段多样化互联网视频的大规模数据集与基准测试平台,用于评估和改进摄像机运动理解。该数据集由专家通过多阶段质量控制流程标注,包含与电影摄影师合作设计的摄像机运动基本类型分类体系。研究表明,如“跟拍”等运动需理解场景内容(如移动主体),而初学者常混淆内参变化(如变焦)与外参变化(如前移),但经教程培训后可显著提升识别准确率。我们在CameraBench上评估了结构光重建(SfM)与视频-语言模型(VLMs),发现SfM模型难以捕捉依赖场景内容的语义运动,而VLMs在精确轨迹估计方面表现不佳。随后,我们基于CameraBench微调生成式VLM,实现语义与几何运动的联合建模,并展示了其在运动增强字幕、视频问答和图文检索中的应用。我们希望该分类体系、基准与教程能推动通用视频摄像机运动理解的发展。
原文摘要 · Abstract (English)
We introduce CameraBench, a large-scale dataset and benchmark designed to assess and improve camera motion understanding. CameraBench consists of ~3,000 diverse internet videos, annotated by experts through a rigorous multi-stage quality control process. One of our contributions is a taxonomy of camera motion primitives, designed in collaboration with cinematographers. We find, for example, that some motions like "follow" (or tracking) require understanding scene content like moving subjects. We conduct a large-scale human study to quantify human annotation performance, revealing that domain expertise and tutorial-based training can significantly enhance accuracy. For example, a novice may confuse zoom-in (a change of intrinsics) with translating forward (a change of extrinsics), but can be trained to differentiate the two. Using CameraBench, we evaluate Structure-from-Motion (SfM) and Video-Language Models (VLMs), finding that SfM models struggle to capture semantic primitives that depend on scene content, while VLMs struggle to capture geometric primitives that require precise estimation of trajectories. We then fine-tune a generative VLM on CameraBench to achieve the best of both worlds and showcase its applications, including motion-augmented captioning, video question answering, and video-text retrieval. We hope our taxonomy, benchmark, and tutorials will drive future efforts towards the ultimate goal of understanding camera motions in any video.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。