arXiv:2606.29461cs.CV2026-06中稿 · ECCV

仅用8张相位图训练出能泛化到复杂物体的次表面散射模型

From Phase to Phenomenon: Self-Supervised Learning of Subsurface Scattering with Minimal Phase-shift Inputs

论文配图:From Phase to Phenomenon: Self-Supervised Learning of Subsurface Scattering with Minimal Phase-shift Inputs
图 1 · 摘自论文原文
  • 用八张高频相位图和立体投影-相机系统实现自监督预训练
  • 在未知几何与材质下仍可高保真重建,比以往方法少用数个数量级图像
  • 适用于材质编辑、光照重演等下游任务,特别适合缺乏标注数据场景

我们提出一种自监督预训练框架,仅需每视角8张高频相位移条纹投影(PSP)图像,即可学习次表面散射(SSS)光传输表示。该方法利用立体投影-相机系统,在多视图、多物体设置中对编码器进行预训练。我们设计了专用于PSP-SSS数据的增强策略,证明其显著优于标准ImageNet风格增强。预训练编码器学习到可泛化的SSS表征,能有效迁移到下游任务,如空间变化光照重演和基于kNN的表征评估。结合专用损失函数训练的解码器,可重建密集的散射足迹响应,尤其在各向异性足迹上提升明显。尽管每视角仅使用8张输入图像,该方法仍能泛化至具有复杂几何和材质的未见物体,实现高保真重建,所需图像数量比先前方法减少数个数量级。

原文摘要 · Abstract (English)

We propose a self-supervised pretraining framework for learning sub-surface scattering (SSS) light transport representations from minimal input. Our method leverages a stereo projector-camera setup that captures only eight high-frequency phase-shift profilometry (PSP) images per view to pretrain an encoder in a multi-view, multi-object setting. We introduce a tailored augmentation strategy for PSP-based SSS data, and show that it significantly outperforms standard ImageNet-style augmentations for SSL pretraining. The pretrained encoder learns generalizable SSS representations that transfer effectively to downstream tasks, including spatially varying relighting and representation evaluation using a kNN classifier. Combined with a decoder, the model reconstructs dense scattering footprint responses, trained using a dedicated cost function that improves accuracy, particularly for anisotropic footprints. Despite using only eight input images per view, our approach generalizes to unseen objects with complex geometry and material properties, achieving high-fidelity reconstructions while requiring orders of magnitude fewer images than prior methods.

次表面散射自监督学习三维重建极少量数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。