动态分配生成模型与资源,降低移动端AI内容生成延迟。
Joint Model Assignment and Resource Allocation for Cost-Effective Mobile Generative Services
- 按提示类别动态选型模型并分配资源
- 生成质量提升4.7%,响应延迟降低39.1%
- 适合边缘计算场景下AIGC服务优化
人工智能生成内容(AIGC)服务可高效满足用户的内容创作需求,但其高计算需求给移动用户规模化支持带来挑战。本文设计了一种边缘赋能的AIGC服务部署系统,将生成模型的计算任务合理分配至边缘服务器,以提升用户体验并降低内容生成延迟。当边缘服务器接收到用户请求的提示后,会根据提示类别特征动态选择合适模型并分配计算资源。关键在于提出一种概率模型分配方法,基于类别标签估计各提示生成内容的质量得分;随后引入启发式算法,根据各生成模型接收到的任务请求自适应配置生成步数与资源分配。仿真结果表明,该系统相较基准方案可使生成内容质量提升最高达4.7%,响应延迟降低最高达39.1%。
原文摘要 · Abstract (English)
Artificial Intelligence Generated Content (AIGC) services can efficiently satisfy user-specified content creation demands, but the high computational requirements pose various challenges to supporting mobile users at scale. In this paper, we present our design of an edge-enabled AIGC service provisioning system to properly assign computing tasks of generative models to edge servers, thereby improving overall user experience and reducing content generation latency. Specifically, once the edge server receives user requested task prompts, it dynamically assigns appropriate models and allocates computing resources based on features of each category of prompts. The generated contents are then delivered to users. The key to this system is a proposed probabilistic model assignment approach, which estimates the quality score of generated contents for each prompt based on category labels. Next, we introduce a heuristic algorithm that enables adaptive configuration of both generation steps and resource allocation, according to the various task requests received by each generative model on the edge.Simulation results demonstrate that the designed system can effectively enhance the quality of generated content by up to 4.7% while reducing response delay by up to 39.1% compared to benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。