arXiv:2506.23150cs.CV2025-06AAAI被引 2

通过分布对齐提升单图生成3D的多视角一致性,速度更快、效果更稳。

AlignCVC: Aligning Cross-View Consistency for Single-Image-to-3D Generation

论文配图:AlignCVC: Aligning Cross-View Consistency for Single-Image-to-3D Generation
图 1 · 摘自论文原文
  • 用分布对齐替代传统回归损失,从根源上改善多视角一致性
  • 仅需4步推理,生成质量与重建性能显著提升
  • 可插拔适配多种生成与重建模型,适合快速部署

单图生成3D模型通常采用串行生成与重建流程。然而,预训练生成模型合成的中间多视角图像常缺乏跨视角一致性(CVC),严重降低3D重建性能。尽管近期方法尝试通过将重建结果反馈至多视角生成器来优化CVC,但受限于噪声大且不稳定的重建输出,改进效果有限。本文提出AlignCVC,一种通过分布对齐重构单图生成3D范式的新框架,而非依赖严格回归损失。核心洞察是将生成与重建的多视角分布共同对齐至真实多视角分布,建立提升CVC的理论基础。观察到生成图像存在弱CVC,而重建图像因显式渲染表现出强CVC,因此提出软-硬对齐策略,为生成与重建模型设定不同目标。该方法不仅提升生成质量,还将推理速度大幅加速至仅需4步。作为即插即用范式,AlignCVC可无缝集成各类多视角生成与3D重建模型。大量实验验证了其在单图生成3D任务中的有效性与高效性。

原文摘要 · Abstract (English)

Single-image-to-3D models typically follow a sequential generation and reconstruction workflow. However, intermediate multi-view images synthesized by pre-trained generation models often lack cross-view consistency (CVC), significantly degrading 3D reconstruction performance. While recent methods attempt to refine CVC by feeding reconstruction results back into the multi-view generator, these approaches struggle with noisy and unstable reconstruction outputs that limit effective CVC improvement. We introduce AlignCVC, a novel framework that fundamentally re-frames single-image-to-3D generation through distribution alignment rather than relying on strict regression losses. Our key insight is to align both generated and reconstructed multi-view distributions toward the ground-truth multi-view distribution, establishing a principled foundation for improved CVC. Observing that generated images exhibit weak CVC while reconstructed images display strong CVC due to explicit rendering, we propose a soft-hard alignment strategy with distinct objectives for generation and reconstruction models. This approach not only enhances generation quality but also dramatically accelerates inference to as few as 4 steps. As a plug-and-play paradigm, our method, namely AlignCVC, seamlessly integrates various multi-view generation models with 3D reconstruction models. Extensive experiments demonstrate the effectiveness and efficiency of AlignCVC for single-image-to-3D generation.

3D生成多视角一致分布对齐快速推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。