arXiv:2505.21483cs.CV2025-05NeurIPS被引 8

实现2D/3D场景中光照阴影一致的高效物体合成

MV-CoLight: Efficient Object Compositing with Consistent Lighting and Shadow Generation

  • 两阶段架构直接建模光照与阴影,避免迭代优化偏差
  • 在标准基准和真实场景中均达到当前最佳融合效果
  • 适用于增强现实与具身智能,支持复杂光照条件

物体合成在增强现实(AR)和具身智能应用中具有巨大潜力。现有方法多集中于单图场景或内在分解技术,在多视角一致性、复杂场景及多样光照条件下存在挑战。近期基于逆渲染的方法(如3D高斯与扩散模型)虽提升了一致性,但受限于可扩展性、数据需求大或每场景重建时间长。为拓展适用性,我们提出MV-CoLight,一种面向2D图像与3D场景的光照一致物体合成两阶段框架。其创新的前馈架构直接建模光照与阴影,规避了扩散方法的迭代偏差。采用希尔伯特曲线映射,实现2D输入与3D高斯场景表示的无缝对齐。为进一步支持训练与评估,我们构建了一个大规模3D合成数据集。实验表明,该框架在标准基准及自建数据集上均达到当前最优融合效果,并在自然拍摄的真实场景中展现出强鲁棒性与广泛泛化能力。

原文摘要 · Abstract (English)

Object compositing offers significant promise for augmented reality (AR) and embodied intelligence applications. Existing approaches predominantly focus on single-image scenarios or intrinsic decomposition techniques, facing challenges with multi-view consistency, complex scenes, and diverse lighting conditions. Recent inverse rendering advancements, such as 3D Gaussian and diffusion-based methods, have enhanced consistency but are limited by scalability, heavy data requirements, or prolonged reconstruction time per scene. To broaden its applicability, we introduce MV-CoLight, a two-stage framework for illumination-consistent object compositing in both 2D images and 3D scenes. Our novel feed-forward architecture models lighting and shadows directly, avoiding the iterative biases of diffusion-based methods. We employ a Hilbert curve-based mapping to align 2D image inputs with 3D Gaussian scene representations seamlessly. To facilitate training and evaluation, we further introduce a large-scale 3D compositing dataset. Experiments demonstrate state-of-the-art harmonized results across standard benchmarks and our dataset, as well as casually captured real-world scenes demonstrate the framework's robustness and wide generalization.

物体合成光照一致3D高斯AR应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。