用相机位姿增强视频理解,让模型像人一样感知空间关系。
Cambrian-P: Pose-Grounded Video Understanding

- 引入可学习的相机位姿令牌和回归头,让模型感知每帧的空间位置。
- 在空间推理任务上提升4.5-6.5%,跨8个基准表现优异。
- 无需真实标注,伪标注位姿也能提升通用视频问答能力。
相机位姿至关重要。每个视角的位置与朝向共同定义了一个共享的空间坐标系,将视频帧间的观测关联起来。然而,当前多模态大模型(MLLM)处理视频时,将帧视为孤立的2D快照,忽略了这一关键信号。我们重新审视位姿作为轻量级监督信号的作用,提出Cambrian-P:一个在帧级别引入可学习相机令牌并配备位姿回归头的视频多模态大模型。通过精心设计的采样策略,该模型在空间推理基准VSI-Bench上取得4.5%-6.5%的显著提升,并在八个额外的空间与通用视频问答基准上实现良好泛化;同时,作为副产品,在ScanNet数据集上达到流式位姿估计的最新水平。令人意外的是,使用真实世界视频中的伪标注位姿进行训练,进一步提升了通用视频问答性能,表明位姿对超越空间推理也有效。这些结果表明,相机位姿是视频模型理解物理世界的基本信号。
原文摘要 · Abstract (English)
Camera pose matters. The position and orientation of each viewpoint define a shared spatial coordinate frame that relates observations across video frames. Yet this signal is largely absent from multimodal LLMs (MLLMs) for video understanding, which process frames as isolated 2D snapshots, instead of the persistent scene humans perceive. We revisit pose as a lightweight supervisory signal and introduce Cambrian-P, a video MLLM augmented with per-frame learnable camera tokens and a pose regression head. With a carefully designed sampling scheme, the model achieves substantial gains of 4.5-6.5% on spatial reasoning benchmarks such as VSI-Bench, generalizes across eight additional spatial and general video QA benchmarks, and, as a byproduct, achieves state of the art streaming pose estimation on ScanNet. Surprisingly, training on pseudo-annotated poses from in-the-wild video further improves general video QA benchmarks, showing pose helps beyond spatial reasoning. Together, these results position camera pose as a fundamental signal for video models that reason about the physical world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。