arXiv:2510.08673cs.CV2025-10中稿 · ICLR被引 17

让模型像摄影师一样思考,统一理解与生成任意视角画面。

Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation

  • 将相机参数视为语言,实现跨视角空间推理。
  • 在400万组视觉-语言-相机数据上训练,支持灵活场景生成。
  • 适合研究多视角理解、虚拟摄影或智能创作的开发者。

以相机为中心的理解与生成是空间智能的核心,但通常被孤立研究。我们提出Puffin,一种统一的相机中心多模态模型,通过扩展相机维度增强空间感知能力。Puffin结合语言回归与基于扩散的生成,可从任意视角解析和生成场景。为弥合相机与视觉-语言之间的模态差距,我们提出新范式:将相机视为语言,引导模型在几何上下文中对齐空间视觉线索与摄影术语。Puffin在包含400万组视觉-语言-相机三元组的大规模数据集Puffin-4M上训练,融合全局相机参数与像素级相机图,实现灵活可靠的时空生成。实验表明,其在相机中心生成与理解任务中优于专用模型。经指令微调后,Puffin可泛化至多样跨视角任务,如空间想象、世界探索与摄影指导。代码、模型、数据集流水线及基准将公开,以推动多模态空间智能研究。

原文摘要 · Abstract (English)

Camera-centric understanding and generation are two cornerstones of spatial intelligence, yet they are typically studied in isolation. We present Puffin, a unified camera-centric multimodal model that extends spatial awareness along the camera dimension. Puffin integrates language regression and diffusion-based generation to interpret and create scenes from arbitrary viewpoints. To bridge the modality gap between cameras and vision-language, we introduce a novel paradigm that treats camera as language, enabling thinking with camera. This guides the model to align spatially grounded visual cues with photographic terminology while reasoning across geometric context. Puffin is trained on Puffin-4M, a large-scale dataset of 4 million vision-language-camera triplets. We incorporate both global camera parameters and pixel-wise camera maps, yielding flexible and reliable spatial generation. Experiments demonstrate Puffin superior performance over specialized models for camera-centric generation and understanding. With instruction tuning, Puffin generalizes to diverse cross-view tasks such as spatial imagination, world exploration, and photography guidance. We will release the code, models, dataset pipeline, and benchmark to advance multimodal spatial intelligence research.

多模态相机视角扩散模型空间理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。