用像素级信息增益融合3D重建与图像生成,提升大视角下的视觉一致性。
FaithFusion: Harmonizing Reconstruction and Generation via Pixel-wise Information Gain
- 基于像素级期望信息增益统一控制生成与重建过程。
- 在6米车道偏移下仍保持FID 107.47,优于现有方法。
- 无需额外条件或结构修改,可直接集成到现有系统中。
在可控驾驶场景重建与3D场景生成中,维持几何保真度的同时,在大幅视角变化下合成视觉合理外观至关重要。然而,几何驱动的3DGS与外观驱动的扩散模型有效融合面临固有挑战,因缺乏像素级、3D一致的编辑标准,常导致过度修复和几何漂移。为此,我们提出 extbf{FaithFusion},一种基于像素级期望信息增益(EIG)驱动的3DGS-扩散模型融合框架。EIG作为统一策略,引导扩散模型以空间先验方式优化高不确定性区域,其像素级权重将编辑结果反向提炼回3DGS。该即插即用系统无需额外先验条件或结构修改。在Waymo数据集上的大量实验表明,本方法在NTA-IoU、NTL-IoU和FID指标上达到当前最优表现,即使在6米车道偏移下仍保持FID 107.47。代码已公开于https://github.com/wangyuanbiubiubiu/FaithFusion。
原文摘要 · Abstract (English)
In controllable driving-scene reconstruction and 3D scene generation, maintaining geometric fidelity while synthesizing visually plausible appearance under large viewpoint shifts is crucial. However, effective fusion of geometry-based 3DGS and appearance-driven diffusion models faces inherent challenges, as the absence of pixel-wise, 3D-consistent editing criteria often leads to over-restoration and geometric drift. To address these issues, we introduce \textbf{FaithFusion}, a 3DGS-diffusion fusion framework driven by pixel-wise Expected Information Gain (EIG). EIG acts as a unified policy for coherent spatio-temporal synthesis: it guides diffusion as a spatial prior to refine high-uncertainty regions, while its pixel-level weighting distills the edits back into 3DGS. The resulting plug-and-play system is free from extra prior conditions and structural modifications.Extensive experiments on the Waymo dataset demonstrate that our approach attains SOTA performance across NTA-IoU, NTL-IoU, and FID, maintaining an FID of 107.47 even at 6 meters lane shift. Our code is available at https://github.com/wangyuanbiubiubiu/FaithFusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。