arXiv:2511.07222cs.CV2025-11中稿 · ICLR被引 17

用多视角图统一建模3D场景理解与生成,验证生成助理解。

Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images

  • 分纹理、几何模块协同建模,实现理解与生成互促。
  • 在VSI-Bench上达55.4分,超越专用3D模型。
  • 适合做3D生成与理解融合研究的开发者参考。

本文提出Omni-View,将统一的多模态理解与生成扩展至基于多视角图像的3D场景,探索“生成促进理解”的原理。该模型由理解模块、纹理模块和几何模块组成,联合建模场景理解、新视角合成与几何估计,实现3D场景理解与生成任务间的协同交互。通过设计,其纹理模块利用时空建模能力进行外观合成,几何模块提供显式几何约束,从而增强对3D场景的整体理解。采用两阶段训练策略,Omni-View在VSI-Bench基准上取得55.4的领先分数,优于现有专用3D理解模型,同时在新视角合成与3D场景生成任务中表现强劲。代码与预训练模型已开源:https://github.com/AIDC-AI/Omni-View。

原文摘要 · Abstract (English)

This paper presents Omni-View, which extends the unified multimodal understanding and generation to 3D scenes based on multiview images, exploring the principle that "generation facilitates understanding". Consisting of understanding model, texture module, and geometry module, Omni-View jointly models scene understanding, novel view synthesis, and geometry estimation, enabling synergistic interaction between 3D scene understanding and generation tasks. By design, it leverages the spatiotemporal modeling capabilities of its texture module responsible for appearance synthesis, alongside the explicit geometric constraints provided by its dedicated geometry module, thereby enriching the model's holistic understanding of 3D scenes. Trained with a two-stage strategy, Omni-View achieves a state-of-the-art score of 55.4 on the VSI-Bench benchmark, outperforming existing specialized 3D understanding models, while simultaneously delivering strong performance in both novel view synthesis and 3D scene generation. The code and pretraiend models are open-sourced at https://github.com/AIDC-AI/Omni-View.

3D理解多视图生成统一建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。