用统一模型生成高保真材质,一次搞定文本、图像转材质。
MatPedia: A Universal Generative Foundation for High-Fidelity Material Synthesis
- 将材质的图像与物理属性联合编码为两个相关潜变量。
- 在1024×1024分辨率下生成效果超越现有方法,质量与多样性更优。
- 适合影视、游戏等领域需要快速生成真实感材质的开发者。
基于物理的渲染(PBR)材质是实现逼真图形的基础,但其创建过程繁琐且需专业技能。尽管生成模型已推动材质合成发展,现有方法缺乏统一表示来连接自然图像外观与PBR属性,导致任务专用流程分散,无法利用大规模RGB图像数据。我们提出MatPedia,一个基于新型联合RGB-PBR表示的基座模型,将材质紧凑编码为两个相互依赖的潜变量:一个用于RGB外观,另一个用于四个编码互补物理属性的PBR贴图。通过将其建模为五帧序列并采用视频扩散架构,MatPedia自然捕捉两者关联,并从RGB生成模型中迁移视觉先验。该联合表示使单一架构可统一处理文本到材质生成、图像到材质生成及内在分解等多种任务。在包含PBR数据集与大规模RGB图像的混合语料库MatHybrid-410K上训练,MatPedia实现原生1024×1024合成,显著优于现有方法在质量和多样性上的表现。
原文摘要 · Abstract (English)
Physically-based rendering (PBR) materials are fundamental to photorealistic graphics, yet their creation remains labor-intensive and requires specialized expertise. While generative models have advanced material synthesis, existing methods lack a unified representation bridging natural image appearance and PBR properties, leading to fragmented task-specific pipelines and inability to leverage large-scale RGB image data. We present MatPedia, a foundation model built upon a novel joint RGB-PBR representation that compactly encodes materials into two interdependent latents: one for RGB appearance and one for the four PBR maps encoding complementary physical properties. By formulating them as a 5-frame sequence and employing video diffusion architectures, MatPedia naturally captures their correlations while transferring visual priors from RGB generation models. This joint representation enables a unified framework handling multiple material tasks--text-to-material generation, image-to-material generation, and intrinsic decomposition--within a single architecture. Trained on MatHybrid-410K, a mixed corpus combining PBR datasets with large-scale RGB images, MatPedia achieves native $1024\times1024$ synthesis that substantially surpasses existing approaches in both quality and diversity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。