arXiv:2605.15660cs.CV2026-05被引 5

无需文本或额外网络,仅用图像实现高质量材质迁移。

MaTe: Images Are All You Need for Material Transfer via Diffusion Transformer

论文配图:MaTe: Images Are All You Need for Material Transfer via Diffusion Transformer
图 1 · 摘自论文原文
  • 在共享潜在空间中通过多模态注意力融合输入图像
  • 零样本训练自由,生成质量超越现有方法
  • 适合追求高效、简洁材质迁移的开发者与研究者

基于扩散的方法进行材质迁移通常依赖文本引导或复杂架构中的辅助网络,面临文本依赖、计算开销大和特征错位等问题。为此,我们提出MaTe,一种简化版扩散框架,摒弃文本引导与参考网络。MaTe在令牌级别整合输入图像,通过共享潜在空间中的多模态注意力实现统一处理。该设计无需额外适配器、ControlNet、反演采样或模型微调。大量实验表明,MaTe在零样本、无训练范式下实现高质量材质生成,视觉质量与效率均优于当前最优方法,同时保持精确细节对齐,显著降低推理门槛。

原文摘要 · Abstract (English)

Recent diffusion-based methods for material transfer rely on image fine-tuning or complex architectures with assistive networks, but face challenges including text dependency, extra computational costs, and feature misalignment. To address these limitations, we propose MaTe, a streamlined diffusion framework that eliminates textual guidance and reference networks. MaTe integrates input images at the token level, enabling unified processing via multi-modal attention in a shared latent space. This design removes the need for additional adapters, ControlNet, inversion sampling, or model fine-tuning. Extensive experiments demonstrate that MaTe achieves high-quality material generation under a zero-shot, training-free paradigm. It outperforms state-of-the-art methods in both visual quality and efficiency while preserving precise detail alignment, significantly simplifying inference prerequisites.

材质迁移扩散模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。