arXiv:2605.25975cs.GRcs.CV2026-05被引 1

用少量视角图生成可随意改变光照的3D高斯模型,速度提升25倍。

F-RNG: Feed-Forward Relightable Neural Gaussians

论文配图:F-RNG: Feed-Forward Relightable Neural Gaussians
图 1 · 摘自论文原文
  • 基于预训练模型和先验知识,直接从稀疏视角生成可重光照的3D高斯
  • 相比现有方法速度提升25倍,质量提高2.0 dB以上
  • 无需微调,可自动适配未来更优的基础模型

从真实物体捕获可重光照的3D资产是广泛研究的问题。基于3D高斯溅射(3DGS)的逐场景优化方法虽支持重光照,但需密集输入视角,且过拟合导致泛化困难。相比之下,通用前馈模型可直接从稀疏视角重建高斯,但结果光照固定,难以重光照。本文提出F-RNG,一种前馈框架,能从稀疏视图直接生成可重光照的3DGS资产。为避免从头训练的巨大开销,F-RNG在现有大型重建模型(LRM)基础上,结合固有分解模型(IDM)的先验。首先引入潜变量插值的细粒度几何合成,增强几何表示;其次提出先验引导的可重光照外观蒸馏,融合IDM先验提取可重光照神经表示;最后通过通用神经渲染器实现灵活高保真重光照。F-RNG无需重新训练或微调底层LRM,可自动受益于未来更优的LRM与IDM。仅需小规模网络,即可在低成本下完成训练,避免在不同光照下重复推理大模型。相比最先进的基于LRM的重光照方法,F-RNG实现约25倍加速,且质量提升约+2.0 dB。

原文摘要 · Abstract (English)

Capturing relightable 3D assets from real-world objects is a widely researched problem. Several per-scene optimization-based methods, based on 3D Gaussian splatting (3DGS), support relighting; however, they usually require dense input views, and their overfitting nature makes it difficult to generalize across scenes. Unlike per-scene optimization methods, generalized feed-forward models can directly reconstruct Gaussians from sparse input views. However, the resulting assets have baked-in illumination and cannot be easily used for relighting. In this paper, we present F-RNG, a feed-forward framework that directly generates relightable 3DGS assets from sparse-view inputs. Training such a model from scratch can require massive data and computing resources, and it is especially challenging to generate relightable assets in a feed-forward manner with acceptable cost. We develop F-RNG upon an existing large reconstruction model (LRM) to extract relightable representations, while also utilizing priors from an intrinsic decomposition model (IDM). Specifically, we first introduce a latent-interpolated fine-grained geometry synthesis to enhance the LRM's geometry representation. Second, we propose a prior-guided relightable appearance distillation to extract relightable neural representations by incorporating IDM priors. Finally, a universal neural renderer enables flexible and high-fidelity relighting. F-RNG requires neither re-training nor fine-tuning of the underlying LRMs, thus can automatically benefit from better LRMs and IDMs in the future. With only small networks that can be trained with affordable data and computational resources, F-RNG avoids the repetitive inference of large models under different light conditions. By comparison to the state-of-the-art LRM-based relighting method, F-RNG achieves ~25x faster relighting, as well as superior quality (~+2.0 dB).

3D高斯重光照前馈模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。