用双向网格预测统一多视角图像色彩,提升3D重建一致性。
CHROMA: Consistent Harmonization of Multi-View Appearance via Bilateral Grid Prediction
- 通过前馈模型预测自适应双向网格校正光照差异。
- 单次处理数百帧,训练时间几乎不变且跨场景泛化强。
- 无需配对数据,结合3D基础模型实现真实世界鲁棒性。
现代相机管线在设备端执行大量处理,如曝光调整、白平衡和色彩校正,虽各自有益,却常导致多视角间出现光度不一致,破坏多视角一致性并损害新视角合成质量。已有方法通过联合优化场景特定表示与每图像外观嵌入来解决此问题,但计算复杂度高且训练慢。本文提出一种可泛化的前馈方法,通过预测空间自适应的双向网格,以多视图一致方式校正光度差异。模型可单步处理数百帧,实现高效大规模调和,并无缝集成至下游3D重建模型,无需场景特定重训练即可实现跨场景泛化。为应对缺乏成对数据的问题,采用混合自监督渲染损失,利用3D基础模型提升对真实世界变化的泛化能力。大量实验表明,本方法在重建质量上优于或匹配现有带外观建模的场景特定优化方法,且对基线3D模型训练时间影响极小。
原文摘要 · Abstract (English)
Modern camera pipelines apply extensive on-device processing, such as exposure adjustment, white balance, and color correction, which, while beneficial individually, often introduce photometric inconsistencies across views. These appearance variations violate multi-view consistency and degrade novel view synthesis. Joint optimization of scene-specific representations and per-image appearance embeddings has been proposed to address this issue, but with increased computational complexity and slower training. In this work, we propose a generalizable, feed-forward approach that predicts spatially adaptive bilateral grids to correct photometric variations in a multi-view consistent manner. Our model processes hundreds of frames in a single step, enabling efficient large-scale harmonization, and seamlessly integrates into downstream 3D reconstruction models, providing cross-scene generalization without requiring scene-specific retraining. To overcome the lack of paired data, we employ a hybrid self-supervised rendering loss leveraging 3D foundation models, improving generalization to real-world variations. Extensive experiments show that our approach outperforms or matches the reconstruction quality of existing scene-specific optimization methods with appearance modeling, without significantly affecting the training time of baseline 3D models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。