单图重光照与三维重建联合优化,提升物理一致性。
GeoRelight: Learning Joint Geometrical Relighting and Reconstruction with Flexible Multi-Modal Diffusion Transformers

- 采用多模态扩散变压器统一建模几何与光照。
- 在真实数据上实现比串行方法更优的重光照效果。
- 适合需要高保真人像重光照的应用场景。
从单张图像进行重光照是一项极具吸引力但病态的任务,因为二维图像模糊地混合了三维几何、固有外观和光照信息。现有方法要么采用串行流水线,易积累误差;要么未在重光照中显式利用三维几何,限制了物理一致性。由于重光照与三维几何估计相互促进,我们提出统一的多模态扩散变压器(DiT)——GeoRelight,联合求解二者。关键技术贡献包括:无畸变的3D表示iNOD(isotropic NDC-Orthographic Depth),兼容潜在扩散模型;以及结合合成数据与自动标注真实数据的混合训练策略。通过联合求解几何与重光照,GeoRelight在性能上优于串行模型及忽略几何的先前系统。
原文摘要 · Abstract (English)
Relighting a person from a single photo is an attractive but ill-posed task, as a 2D image ambiguously entangles 3D geometry, intrinsic appearance, and illumination. Current methods either use sequential pipelines that suffer from error accumulation, or they do not explicitly leverage 3D geometry during relighting, which limits physical consistency. Since relighting and estimation of 3D geometry are mutually beneficial tasks, we propose a unified Multi-Modal Diffusion Transformer (DiT) that jointly solves for both: GeoRelight. We make this possible through two key technical contributions: isotropic NDC-Orthographic Depth (iNOD), a distortion-free 3D representation compatible with latent diffusion models; and a strategic mixed-data training method that combines synthetic and auto-labeled real data. By solving geometry and relighting jointly, GeoRelight achieves better performance than both sequential models and previous systems that ignored geometry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。