让手机快速预览,云端精细生成,提升图像生成效率。
DiffusionX: Efficient Edge-Cloud Collaborative Image Generation with Multi-Round Prompt Evolution
- 手机端轻量模型实时出图,云端高精度模型最终优化。
- 相比Stable Diffusion v1.5,平均生成时间减少15.8%。
- 适合对延迟敏感、资源受限的移动端图像生成场景。
扩散模型的进展推动了图像生成的显著进步,但生成过程仍计算密集,用户常需多次迭代优化提示词,进一步增加延迟并加重云资源负担。为此,我们提出DiffusionX,一种面向多轮提示生成的云-边协同框架。该系统中,轻量级设备端扩散模型快速生成预览图并与用户交互,待提示词确定后,由高性能云模型完成最终精修。我们还引入噪声级别预测器,动态平衡计算负载,优化延迟与云工作量之间的权衡。实验表明,DiffusionX相较于Stable Diffusion v1.5平均生成时间降低15.8%,同时保持相当的图像质量;其仅比Tiny-SD慢0.9%,但图像质量显著提升,证明了其高效性与可扩展性,且开销极小。
原文摘要 · Abstract (English)
Recent advances in diffusion models have driven remarkable progress in image generation. However, the generation process remains computationally intensive, and users often need to iteratively refine prompts to achieve the desired results, further increasing latency and placing a heavy burden on cloud resources. To address this challenge, we propose DiffusionX, a cloud-edge collaborative framework for efficient multi-round, prompt-based generation. In this system, a lightweight on-device diffusion model interacts with users by rapidly producing preview images, while a high-capacity cloud model performs final refinements after the prompt is finalized. We further introduce a noise level predictor that dynamically balances the computation load, optimizing the trade-off between latency and cloud workload. Experiments show that DiffusionX reduces average generation time by 15.8% compared with Stable Diffusion v1.5, while maintaining comparable image quality. Moreover, it is only 0.9% slower than Tiny-SD with significantly improved image quality, thereby demonstrating efficiency and scalability with minimal overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。