arXiv:2512.00493cs.CV2025-12

零样本单图生成3D场景,保持物体布局与相机一致性

CC-FMO: Camera-Conditioned Zero-Shot Single Image to 3D Scene Generation with Foundation Model Orchestration

论文配图:CC-FMO: Camera-Conditioned Zero-Shot Single Image to 3D Scene Generation with Foundation Model Orchestration
图 1 · 摘自论文原文
  • 融合语义向量与结构化隐变量生成高质量物体
  • 通过相机条件解尺度算法提升场景空间一致性
  • 无需微调,适合快速构建真实感3D场景

从单张图像生成高质量3D场景对AR/VR和具身AI至关重要。早期方法因依赖小规模定制数据集上的专用模型而泛化能力差。尽管大规模3D基础模型显著提升了实例级生成效果,但场景级生成仍面临挑战,主要受限于物体位姿估计不准和空间不一致。为此,本文提出CC-FMO,一种零样本、相机条件化的单图到3D场景生成框架,能同时遵循输入图像中的物体布局并保持实例保真度。CC-FMO采用混合实例生成器,结合语义感知的向量集表示与细节丰富的结构化隐变量表示,生成既语义合理又高保真的物体几何。此外,通过简单有效的相机条件化尺度求解算法,将基础位姿估计模型应用于场景生成任务,实现场景级一致性。大量实验表明,CC-FMO始终生成高保真、相机对齐的组合式场景,优于所有现有方法。

原文摘要 · Abstract (English)

High-quality 3D scene generation from a single image is crucial for AR/VR and embodied AI applications. Early approaches struggle to generalize due to reliance on specialized models trained on curated small datasets. While recent advancements in large-scale 3D foundation models have significantly enhanced instance-level generation, coherent scene generation remains a challenge, where performance is limited by inaccurate per-object pose estimations and spatial inconsistency. To this end, this paper introduces CC-FMO, a zero-shot, camera-conditioned pipeline for single-image to 3D scene generation that jointly conforms to the object layout in input image and preserves instance fidelity. CC-FMO employs a hybrid instance generator that combines semantics-aware vector-set representation with detail-rich structured latent representation, yielding object geometries that are both semantically plausible and high-quality. Furthermore, CC-FMO enables the application of foundational pose estimation models in the scene generation task via a simple yet effective camera-conditioned scale-solving algorithm, to enforce scene-level coherence. Extensive experiments demonstrate that CC-FMO consistently generates high-fidelity camera-aligned compositional scenes, outperforming all state-of-the-art methods.

3D生成单图生成基础模型场景重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。