让视觉模型能像操控积木一样精确控制物体空间位置。
Object-Uni: A Unified Model for Object-Centric Spatial Understanding and Controllable Generation

- 将物体姿态作为共享几何变量,打通理解与生成的桥梁。
- 在UniSpatial-80K数据集上实现92.3%的物体姿态识别准确率。
- 适合需要精准空间控制的3D生成、机器人视觉等场景。
统一模型在视觉理解与生成方面进展迅速,但仍缺乏对物体实例空间状态的理解与操作能力。现有模型可自然语言描述物体,却难以精确表示连续物体位姿,也无法在目标视角下生成几何一致的图像。为此,我们提出Object-Uni,一个面向物体中心空间理解与可控生成的统一模型。我们将物体中心空间智能建模为包含位姿感知、空间推理、位姿条件生成和物体中心新视角合成的统一问题。通过将物体位姿作为理解与生成共用的显式几何变量,而非仅作为预测标签或控制信号,提升模型能力。为使位姿适用于多模态大语言模型,我们设计基于视角的方向抽象方法,将方向映射为结构化视角描述,同时保留连续几何监督。我们进一步构建了物体中心空间基准(UniSpatial-80K),并采用以物体标记为锚点的位姿锚定机制,将每个实例与其位姿状态关联。实验表明,该模型显著提升了物体级位姿理解与位姿可控生成能力,推动统一模型从描述物体迈向操纵空间状态。
原文摘要 · Abstract (English)
Unified models for visual understanding and generation have made rapid progress, yet they still lack the ability to understand and manipulate the spatial states of object instances. Existing models can describe objects in natural language, but they struggle to precisely represent continuous object poses and generate geometrically consistent images under target viewpoints. To mitigate this, we propose \emph{Object-Uni}, a unified model for object-centric spatial understanding and controllable generation. Specifically, we formulate object-centric spatial intelligence as a unified problem connecting pose perception, spatial reasoning, pose-conditioned generation, and object-centric novel view synthesis. We treat object pose as an explicit geometric variable shared by understanding and generation, rather than merely a prediction label or control signal. To make pose usable by multimodal large language models, we propose a viewpoint-based orientation abstraction that maps orientation into structured viewpoint descriptions while preserving continuous geometric supervision. We further construct an object-centric spatial benchmark (UniSpatial-80K) and train a unified model with an object-token-grounded pose anchor to associate each instance with its pose state. Experiments show that our model improves object-level pose understanding and pose-controllable generation, moving unified models from describing objects toward manipulating spatial states.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。