arXiv:2502.06798cs.LGcs.DC2025-02被引 2

根据提示词动态调度模型,高效生成高质量图像。

Prompt-Aware Scheduling for Efficient Text-to-Image Inferencing System

  • 按提示词匹配不同精度的模型实例,实现动态调度
  • 高负载下保持图像质量,资源使用在固定预算内
  • 适合需要快速生成且对质量要求高的场景

传统机器学习模型在高负载时通过使用更快但精度较低的模型进行精度缩放。然而,由于生成式文本到图像模型对输入提示敏感,且大模型加载开销导致性能下降,该方法效果不佳。本文提出一种新型文本到图像推理系统,通过在多个运行于不同近似水平的同一模型实例间,最优匹配提示词,实现在高负载和固定资源预算下输出高质量图像。

原文摘要 · Abstract (English)

Traditional ML models utilize controlled approximations during high loads, employing faster, but less accurate models in a process called accuracy scaling. However, this method is less effective for generative text-to-image models due to their sensitivity to input prompts and performance degradation caused by large model loading overheads. This work introduces a novel text-to-image inference system that optimally matches prompts across multiple instances of the same model operating at various approximation levels to deliver high-quality images under high loads and fixed budgets.

文本生成图像生成推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。