用扩散模型高效生成4K级材质贴图,解决高分辨率与像素对齐难题。
HiMat: DiT-based Ultra-High Resolution SVBRDF Generation
- 在高压缩潜空间中生成多张反射图,降低显存与计算开销
- 4K分辨率下保持各材质图间像素级对齐,质量优于现有方法
- 适合需要高保真材质生成的影视、游戏与数字孪生场景
生成超高清空间变化双向反射分布函数(SVBRDF)对于实现逼真的3D内容至关重要,可精准呈现近景渲染所需的微观表面细节。然而,实现4K生成面临两大挑战:(1) 需在全分辨率下合成多张反射图,导致像素预算倍增,带来难以承受的内存与计算成本;(2) 要求在4K分辨率下维持各图间强像素级对齐,而预训练模型通常针对RGB图像域设计,难以适配。本文提出HiMat,一种基于扩散变换器的高效且多样化的4K SVBRDF生成框架。为应对第一项挑战,HiMat通过DC-AE在高压缩潜空间中进行生成,并采用具备线性注意力机制的预训练扩散变换器提升单图效率;为应对第二项挑战,提出CrossStitch——一种轻量级卷积模块,无需全局注意力即可强制跨图一致性。实验表明,与先前方法相比,HiMat在4K分辨率下实现了更高保真度、更优效率、更强结构一致性和更大多样性。该框架还可泛化至如内在分解等其他应用。
原文摘要 · Abstract (English)
Creating ultra-high-resolution spatially varying bidirectional reflectance functions (SVBRDFs) is critical for photorealistic 3D content creation, to faithfully represent fine-scale surface details required for close-up rendering. However, achieving 4K generation faces two key challenges: (1) the need to synthesize multiple reflectance maps at full resolution, which multiplies the pixel budget and imposes prohibitive memory and computational cost, and (2) the requirement to maintain strong pixel-level alignment across maps at 4K, which is particularly difficult when adapting pretrained models designed for the RGB image domain. We introduce HiMat, a diffusion-based framework tailored for efficient and diverse 4K SVBRDF generation. To address the first challenge, HiMat performs generation in a high-compression latent space via DC-AE, and employs a pretrained diffusion transformer with linear attention to improve per-map efficiency. To address the second challenge, we propose CrossStitch, a lightweight convolutional module that enforces cross-map consistency without incurring the cost of global attention. Our experiments show that HiMat achieves high-fidelity 4K SVBRDF generation with superior efficiency, structural consistency, and diversity compared to prior methods. Beyond materials, our framework also generalizes to related applications such as intrinsic decomposition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。