用视频扩散模型生成3D材质,支持文本和几何输入。
VideoMatGen: PBR Materials through Joint Generative Modeling
- 基于视频扩散架构,联合建模多种材质属性。
- 通过自定义变分自编码器压缩多模态材质特征。
- 支持文本+几何输入,输出可直接用于创作工具。
我们提出一种基于视频扩散变压器架构的3D材质生成方法,该方法以输入几何体和文本描述为条件,联合建模多种材质属性(基础颜色、粗糙度、金属度、高度图),生成物理上合理的材质。我们引入一个定制的变分自编码器,将多种材质模态编码到紧凑的潜在空间中,实现在不增加令牌数量的前提下联合生成多个模态。该流程可根据文本提示生成高质量3D材质,兼容常见内容创作工具。
原文摘要 · Abstract (English)
We present a method for generating physically-based materials for 3D shapes based on a video diffusion transformer architecture. Our method is conditioned on input geometry and a text description, and jointly models multiple material properties (base color, roughness, metallicity, height map) to form physically plausible materials. We further introduce a custom variational auto-encoder which encodes multiple material modalities into a compact latent space, which enables joint generation of multiple modalities without increasing the number of tokens. Our pipeline generates high-quality materials for 3D shapes given a text prompt, compatible with common content creation tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。