Nexus通过结构化协同设计,实现高效文本生成图像。
Nexus: Structured Synergy for Efficient Text-to-Image Generation using Rectified Flow Model

- 采用稀疏架构与线性复杂度设计,降低计算开销。
- 在COCO和LAION数据集上达到SDXL级画质,推理效率显著提升。
- 适合边缘部署及高分辨率图像生成场景。
扩散模型与流匹配模型在文本到图像生成中取得显著进展,但高计算量、二次复杂度和大内存占用限制了高分辨率合成与边缘部署。我们提出Nexus,融合稀疏架构、线性复杂度与低比特量化。通过引入MoE前馈层、门控DeltaNet注意力机制以及每专家低比特训练,有效减少计算与内存消耗。三者联合优化使Nexus在保持与主流模型(如SDXL、SD3)相当生成质量的同时,显著提升推理效率。在COCO和LAION数据集上的实验验证了其有效性。
原文摘要 · Abstract (English)
Diffusion and flow matching models have made significant progress in text-to-image generation, yet high computation, quadratic complexity, and large memory footprint hinder high-resolution synthesis and edge deployment. We propose Nexus, which integrates sparse architecture, linear complexity, and low-bit quantization. It combines MoE feed-forward layers, gated DeltaNet attention, and per-expert low-bit training to reduce computation and memory. Their joint optimization allows Nexus to achieve generation quality comparable to mainstream models such as SDXL and SD3 while delivering markedly higher inference efficiency. Experiments on COCO and LAION validate its effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。