用表面光场编码统一建模3D形状与视角依赖外观。
LiTo: Surface Light Field Tokenization
- 从RGB-D图像中采样表面光场,用潜在向量联合表示几何与外观。
- 生成物体在复杂光照下具备镜面高光和菲涅耳反射等真实视觉效果。
- 仅需单张输入图即可生成符合光照材质的3D对象,适合内容生成场景。
我们提出一种3D潜在表示,可联合建模物体几何与视角依赖外观。以往多数工作仅关注3D几何重建或视角无关漫反射外观预测,难以捕捉真实视角依赖效应。我们的方法利用RGB-D图像提供表面光场的采样点,通过将随机子采样的表面光场编码为紧凑的潜在向量集合,使模型在统一的3D潜在空间中学习几何与外观表征。该表示能复现复杂光照下的视角依赖效果,如镜面高光和菲涅耳反射。我们进一步在此表示上训练潜在流匹配模型,使其基于单张输入图像学习分布,从而生成与输入光照和材质一致的3D物体。实验表明,该方法在视觉质量和输入保真度上优于现有技术。
原文摘要 · Abstract (English)
We propose a 3D latent representation that jointly models object geometry and view-dependent appearance. Most prior works focus on either reconstructing 3D geometry or predicting view-independent diffuse appearance, and thus struggle to capture realistic view-dependent effects. Our approach leverages that RGB-depth images provide samples of a surface light field. By encoding random subsamples of this surface light field into a compact set of latent vectors, our model learns to represent both geometry and appearance within a unified 3D latent space. This representation reproduces view-dependent effects such as specular highlights and Fresnel reflections under complex lighting. We further train a latent flow matching model on this representation to learn its distribution conditioned on a single input image, enabling the generation of 3D objects with appearances consistent with the lighting and materials in the input. Experiments show that our approach achieves higher visual quality and better input fidelity than existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。