arXiv:2511.13684cs.CVcs.LG2025-11被引 1

无需训练,用文本指令实现多视角3D场景精准光影重渲染。

Training-Free Multi-View Extension of IC-Light for Textual Position-Aware Scene Relighting

  • 通过大视觉语言模型解析文本指令,结合几何语义信息生成光照先验。
  • 多视角输入下生成高质量、方向一致的重光照图像,提升视觉一致性。
  • 适合需要快速、精准控制3D场景光照的创作者与设计师使用。

我们提出GS-Light,一种高效、文本位置感知的3D场景重光照方法,适用于基于高斯溅射(3DGS)表示的3D场景。该方法在无需训练的前提下,将单输入扩散模型扩展至多视角输入。用户可通过指定光照方向、颜色、强度或参考物体的文本提示,由大视觉语言模型(LVLM)解析为光照先验。结合现成的几何与语义估计器(深度、法向量、语义分割),融合光照先验与视图几何约束,计算出各视角的光照图并生成初始潜在代码。这些精细构建的初始潜码引导扩散模型生成更符合用户期望的重光照结果,尤其在光照方向上表现更准确。将多视角渲染图像与初始潜码输入多视角重光照模型,生成高保真、艺术化重光照图像。最后,利用重光照外观对3DGS场景进行微调,获得完整重光照3D场景。我们在室内外场景上评估了GS-Light,对比了最先进的逐视图重光照、视频重光照及场景编辑方法。通过定量指标(多视角一致性、成像质量、美学评分、语义相似度等)和定性评估(用户研究),结果表明其持续优于基线。代码与资源将在发表后公开。

原文摘要 · Abstract (English)

We introduce GS-Light, an efficient, textual position-aware pipeline for text-guided relighting of 3D scenes represented via Gaussian Splatting (3DGS). GS-Light implements a training-free extension of a single-input diffusion model to handle multi-view inputs. Given a user prompt that may specify lighting direction, color, intensity, or reference objects, we employ a large vision-language model (LVLM) to parse the prompt into lighting priors. Using off-the-shelf estimators for geometry and semantics (depth, surface normals, and semantic segmentation), we fuse these lighting priors with view-geometry constraints to compute illumination maps and generate initial latent codes for each view. These meticulously derived init latents guide the diffusion model to generate relighting outputs that more accurately reflect user expectations, especially in terms of lighting direction. By feeding multi-view rendered images, along with the init latents, into our multi-view relighting model, we produce high-fidelity, artistically relit images. Finally, we fine-tune the 3DGS scene with the relit appearance to obtain a fully relit 3D scene. We evaluate GS-Light on both indoor and outdoor scenes, comparing it to state-of-the-art baselines including per-view relighting, video relighting, and scene editing methods. Using quantitative metrics (multi-view consistency, imaging quality, aesthetic score, semantic similarity, etc.) and qualitative assessment (user studies), GS-Light demonstrates consistent improvements over baselines. Code and assets will be made available upon publication.

3D重光照文本控制多视角扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。