从视频模型中提取可复用的神经材质,实现逼真3D材料生成。
VideoNeuMat: Neural Material Extraction from Generative Video Models
- 通过控制光照与摄像机轨迹,将视频模型转为虚拟光泽仪。
- 仅用17帧视频即可单次推理生成泛化性强的神经材质参数。
- 适合3D渲染、游戏开发等需要高质量材质的领域。
创建用于3D渲染的逼真材质需极高艺术技能。尽管生成式材质模型有潜力,但受限于高质量训练数据不足。近期视频生成模型能轻松生成逼真材质外观,但其知识仍与几何和光照纠缠。我们提出VideoNeuMat,一种两阶段流程,从视频扩散模型中提取可复用的神经材质资产。首先,微调大型视频模型(Wan 2.1 14B)在受控相机与光照轨迹下生成材质样本视频,形成“虚拟光泽仪”,在保留材质真实感的同时学习结构化测量模式。其次,通过微调较小的Wan 1.3B视频骨干网络构建的大重建模型(LRM),从这些视频中重建紧凑的神经材质。仅需17个生成视频帧,该模型即可单次推理预测出可泛化至新视角与光照条件的神经材质参数。生成材质在真实感与多样性上远超有限的合成训练数据,证明了可将互联网规模视频模型中的材质知识成功迁移为独立、可重用的神经3D资产。
原文摘要 · Abstract (English)
Creating photorealistic materials for 3D rendering requires exceptional artistic skill. Generative models for materials could help, but are currently limited by the lack of high-quality training data. While recent video generative models effortlessly produce realistic material appearances, this knowledge remains entangled with geometry and lighting. We present VideoNeuMat, a two-stage pipeline that extracts reusable neural material assets from video diffusion models. First, we finetune a large video model (Wan 2.1 14B) to generate material sample videos under controlled camera and lighting trajectories, effectively creating a "virtual gonioreflectometer" that preserves the model's material realism while learning a structured measurement pattern. Second, we reconstruct compact neural materials from these videos through a Large Reconstruction Model (LRM) finetuned from a smaller Wan 1.3B video backbone. From 17 generated video frames, our LRM performs single-pass inference to predict neural material parameters that generalize to novel viewing and lighting conditions. The resulting materials exhibit realism and diversity far exceeding the limited synthetic training data, demonstrating that material knowledge can be successfully transferred from internet-scale video models into standalone, reusable neural 3D assets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。