arXiv:2411.13787cs.CVcs.LG2024-11ICCV被引 2

智能分配图文生成任务,兼顾质量与成本。

Adaptive Routing of Text-to-Image Generation Requests Between Large Cloud Model and Light-Weight Edge Model

  • 根据提示词特征动态选择云端或边缘模型
  • 多维度评估图像质量,减少对大模型依赖
  • 适合注重成本与效率的AI生成应用

大型文本到图像模型生成效果出色,但需昂贵云服务器部署;轻量级模型可低成本部署于边缘设备,但复杂提示下的生成质量较差。为平衡性能与成本,我们提出路由框架RouteT2I,针对每个用户提示动态选择使用大型云模型或轻量级边缘模型。由于图像质量难以直接衡量,RouteT2I构建多维质量指标,通过比对生成图像与正负文本描述的相似性来评估。该框架识别提示中的关键标记,分析其对质量的影响,并引入帕累托相对优势比较多指标生成质量。基于此比较和预设成本约束,决定将请求分配至边缘或云端。评估表明,RouteT2I显著降低对大型云模型的请求量,同时保持高质量图像生成。

原文摘要 · Abstract (English)

Large text-to-image models demonstrate impressive generation capabilities; however, their substantial size necessitates expensive cloud servers for deployment. Conversely, light-weight models can be deployed on edge devices at lower cost but often with inferior generation quality for complex user prompts. To strike a balance between performance and cost, we propose a routing framework, called RouteT2I, which dynamically selects either the large cloud model or the light-weight edge model for each user prompt. Since generated image quality is challenging to measure and compare directly, RouteT2I establishes multi-dimensional quality metrics, particularly, by evaluating the similarity between the generated images and both positive and negative texts that describe each specific quality metric. RouteT2I then predicts the expected quality of the generated images by identifying key tokens in the prompt and comparing their impact on the quality. RouteT2I further introduces the Pareto relative superiority to compare the multi-metric quality of the generated images. Based on this comparison and predefined cost constraints, RouteT2I allocates prompts to either the edge or the cloud. Evaluation reveals that RouteT2I significantly reduces the number of requesting large cloud model while maintaining high-quality image generation.

图文生成边缘计算智能路由成本优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。