arXiv:2607.03043cs.CV2026-07中稿 · ECCV

提出新任务与数据集,提升模型理解摄像机运动的能力。

Natural Language Camera Movement Understanding

论文配图:Natural Language Camera Movement Understanding
图 1 · 摘自论文原文
  • 构建电影级运动分类体系,设计真实与合成视频的原子级评测基准。
  • 微调后模型在真实与合成视频上分别优于Gemini 3.1 Pro 10%和11%。
  • 揭示现有模型对镜头运动识别的严重缺陷,适合视频生成研究者参考。

理解自然语言中的摄像机运动对于训练和评估视频生成模型至关重要。然而我们发现,现有视觉-语言模型(VLMs)在此任务上表现令人意外地差,常混淆平移与旋转、左右方向以及物体运动与摄像机运动。为此,我们确立“自然语言摄像机运动理解”为独立研究任务。提出两级电影学分类体系,构建包含真实与合成视频的大规模原子级评测基准。同时,整理大规模多源训练集,并通过针对性摄像机运动增强提升性能。微调后的VLM-8B在该基准的真实与合成视频上分别优于Gemini 3.1 Pro 10%和11%。尽管取得进展,与人类表现仍有显著差距,凸显该领域亟需更多研究关注。

原文摘要 · Abstract (English)

Understanding camera movement in natural language is critical for training and evaluating video generation models, among other applications. However, we demonstrate that existing vision-language models (VLMs) fail this task in surprising ways, frequently confusing translation with rotation, left with right, and object movement with camera movement. To address these limitations, we establish natural language camera movement understanding as a standalone research task. We introduce a two-level cinematographic taxonomy and an extensive, atomic benchmark featuring both real and synthetic videos. Furthermore, we curate a large-scale, multi-source training set enhanced by targeted camera movement augmentation. Our fine-tuned VLM-8B outperforms Gemini 3.1 Pro by 10% and 11% on our benchmark's real and synthetic videos, respectively. Despite these gains, a significant gap remains relative to human performance, underscoring the need to promote and facilitate future research on natural language camera movement understanding.

视频理解视觉语言摄像机运动评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。