用图文提示生成高质量材质,能修复扭曲或遮挡的图像。
MaterialPicker: Multi-Modal DiT-Based Material Generation

- 基于扩散Transformer,将材质视为视频帧进行生成
- 支持从局部模糊/倾斜照片生成完整材质,且多样性优于以往方法
- 适合虚拟场景制作、逆向渲染等需要真实材质的应用
高质量材质生成对虚拟环境创作和逆向渲染至关重要。我们提出MaterialPicker,一种基于扩散Transformer(DiT)的多模态材质生成方法,可依据文本提示和/或照片生成高保真材质。该方法能从材质样本的图像裁剪中生成完整材质,即使表面存在扭曲、视角倾斜或部分遮挡,仍可有效重建。我们通过微调预训练的DiT视频生成器实现材质生成,将每个材质图视为视频序列中的一帧。在定量与定性评估中,结果表明该方法生成的材质更丰富多样,且在畸变修正方面显著优于现有工作。
原文摘要 · Abstract (English)
High-quality material generation is key for virtual environment authoring and inverse rendering. We propose MaterialPicker, a multi-modal material generator leveraging a Diffusion Transformer (DiT) architecture, improving and simplifying the creation of high-quality materials from text prompts and/or photographs. Our method can generate a material based on an image crop of a material sample, even if the captured surface is distorted, viewed at an angle or partially occluded, as is often the case in photographs of natural scenes. We further allow the user to specify a text prompt to provide additional guidance for the generation. We finetune a pre-trained DiT-based video generator into a material generator, where each material map is treated as a frame in a video sequence. We evaluate our approach both quantitatively and qualitatively and show that it enables more diverse material generation and better distortion correction than previous work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。