arXiv:2510.08271cs.GRcs.CV2025-10ICCV被引 8

单图生成可光照编辑的3D材质,让虚拟资产更真实。

SViM3D: Stable Video Material Diffusion for Single Image 3D Generation

  • 用视频扩散模型联合生成多视角材质与法线,支持物理渲染。
  • 在多个数据集上实现顶尖的光影重演与新视角合成效果。
  • 适合游戏、影视、AR/VR等需要可编辑3D资产的场景。

我们提出Stable Video Materials 3D(SViM3D),一种仅需单张图像即可预测多视角一致的基于物理的渲染(PBR)材质的框架。近期视频扩散模型已成功用于高效重建3D物体。然而,反照率仍常由简单材质模型表示,或需额外步骤估计以实现光照重演与可控外观编辑。我们扩展了潜在视频扩散模型,使其在显式相机控制下联合输出空间变化的PBR参数与表面法线。该设计使模型能作为神经先验,支持光照重演与3D资产生成。我们引入多种机制提升该病态问题下的生成质量。在多个以物体为中心的数据集上,我们的方法在光照重演与新视角合成任务中达到当前最佳性能。模型具备良好泛化能力,适用于多样化输入,生成可用于AR/VR、电影、游戏等视觉媒体的可光照编辑3D资产。

原文摘要 · Abstract (English)

We present Stable Video Materials 3D (SViM3D), a framework to predict multi-view consistent physically based rendering (PBR) materials, given a single image. Recently, video diffusion models have been successfully used to reconstruct 3D objects from a single image efficiently. However, reflectance is still represented by simple material models or needs to be estimated in additional steps to enable relighting and controlled appearance edits. We extend a latent video diffusion model to output spatially varying PBR parameters and surface normals jointly with each generated view based on explicit camera control. This unique setup allows for relighting and generating a 3D asset using our model as neural prior. We introduce various mechanisms to this pipeline that improve quality in this ill-posed setting. We show state-of-the-art relighting and novel view synthesis performance on multiple object-centric datasets. Our method generalizes to diverse inputs, enabling the generation of relightable 3D assets useful in AR/VR, movies, games and other visual media.

3D生成材质生成扩散模型可光照编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。