arXiv:2509.06068cs.CV2025-09

用消费级显卡训练出高质量文生图模型,成本仅约600美元。

Home-made Diffusion Model from Scratch to Hatch

  • 创新的跨注意力跳接U型变压器,提升图像组合一致性。
  • 四张RTX5090训练1024×1024图像,成本535-620美元。
  • 小模型(343M参数)也能实现直观相机控制等高级功能。

我们提出家用扩散模型(HDM),一种高效且强大的文生图扩散模型,专为在消费级硬件上训练(及推理)优化。HDM在保持极低训练成本(535-620美元)的同时,实现了媲美主流模型的1024×1024生成质量,仅需四张RTX5090 GPU。关键贡献包括:(1) 跨注意力跳接的U型变压器(Cross-U-Transformer, XUT),通过跨注意力机制增强特征融合,显著提升图像组合一致性;(2) 一套完整训练方案,包含TREAD加速、新颖的偏移平方裁剪策略以实现高效任意长宽比训练,以及渐进式分辨率扩展;(3) 实证表明,经过精心设计的小模型(343M参数)可达成高质量生成结果,并涌现出直观相机控制等能力。本工作提供了一种新的模型扩展范式,为资源有限的个人研究者和小型机构实现高质量文生图生成提供了可行路径。

原文摘要 · Abstract (English)

We introduce Home-made Diffusion Model (HDM), an efficient yet powerful text-to-image diffusion model optimized for training (and inferring) on consumer-grade hardware. HDM achieves competitive 1024x1024 generation quality while maintaining a remarkably low training cost of $535-620 using four RTX5090 GPUs, representing a significant reduction in computational requirements compared to traditional approaches. Our key contributions include: (1) Cross-U-Transformer (XUT), a novel U-shape transformer, Cross-U-Transformer (XUT), that employs cross-attention for skip connections, providing superior feature integration that leads to remarkable compositional consistency; (2) a comprehensive training recipe that incorporates TREAD acceleration, a novel shifted square crop strategy for efficient arbitrary aspect-ratio training, and progressive resolution scaling; and (3) an empirical demonstration that smaller models (343M parameters) with carefully crafted architectures can achieve high-quality results and emergent capabilities, such as intuitive camera control. Our work provides an alternative paradigm of scaling, demonstrating a viable path toward democratizing high-quality text-to-image generation for individual researchers and smaller organizations with limited computational resources.

文生图扩散模型小模型低成本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。