arXiv:2602.00181cs.CVcs.AI2026-02被引 8

让模型像导演一样思考镜头运动,提升视频理解准确性。

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning

  • 采用观察-推理-回答框架,强制模型进行结构化空间推理。
  • 在多个数据集上将分类准确率提升至78.4%,VQA达74.5%。
  • 首次用强化学习对齐镜头逻辑,适合影视分析与视频理解研究者。

理解相机运动是视频空间智能的基础。然而,现有多模态模型多将其视为黑箱分类任务,常因依赖表面视觉模式而非几何线索而混淆物理上不同的运动。我们提出CamReasoner,将相机运动理解重构为结构化推理过程,以弥合感知与电影逻辑之间的差距。该方法基于观察-思考-回答(O-T-A)范式,要求模型在显式推理模块中阐述时空观察并推理解释运动模式。为实现此能力,我们构建了包含18,000条SFT推理链和38,000条强化学习反馈样本的大规模推理轨迹套件。据我们所知,这是首个在相机运动理解中使用强化学习进行逻辑对齐的工作,确保运动推断基于结构化视觉推理而非上下文猜测。基于Qwen2.5-VL-7B,CamReasoner-7B将二分类准确率从73.8%提升至78.4%,VQA准确率从60.9%提升至74.5%,在多个基准测试中持续优于主流开源与闭源模型。

原文摘要 · Abstract (English)

Understanding camera dynamics is a fundamental pillar of video spatial intelligence. However, existing multimodal models predominantly treat this task as a black-box classification, often confusing physically distinct motions by relying on superficial visual patterns rather than geometric cues. We present \textbf{CamReasoner}, a framework that reformulates camera movement understanding as a structured inference process to bridge the gap between perception and cinematic logic. Our approach centers on the Observation-Thinking-Answer (O-T-A) paradigm, which compels the model to articulate spatio-temporal observations and reason about motion patterns within an explicit reasoning block. To instill this capability, we construct a Large-scale Inference Trajectory Suite comprising 18k SFT reasoning chains and 38k RL feedback samples. To the best of our knowledge, \textbf{we are the first to employ RL for logical alignment in camera movement understanding}, ensuring motion inferences are grounded in structured visual reasoning rather than contextual guesswork. Built upon Qwen2.5-VL-7B, CamReasoner-7B improves binary classification accuracy from 73.8\% to 78.4\% and VQA accuracy from 60.9\% to 74.5\% over its backbone, consistently outperforming both proprietary and open-source baselines across multiple benchmarks.

视频理解结构推理强化学习相机运动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。