解决移动AIGC网络中显存不足导致的服务失败问题。
A-MADiff: Attention-Guided Multi-Agent DRL with Diffusion Policies for Memory-Aware Task Orchestration in Mobile AIGC Networks

- 多智能体强化学习+扩散策略,协同调度任务
- 显存异构下仍能提升整体奖励达显著水平
- 适合研究边缘计算与AI服务调度的学者
人工智能生成内容(AIGC)服务利用生成式AI(GenAI)模型自动生成多样化内容。移动AIGC网络将GenAI模型部署在边缘位置的AIGC服务提供商(ASPs)上,为移动用户提供低延迟、个性化的AIGC服务。然而,AIGC推理任务通常占用GPU内存直至完成,导致服务端显存耗尽并引发内存溢出错误,而非仅增加延迟。现有AIGC任务调度研究大多忽略了显存可行性约束。为此,我们提出一种协作式多智能体调度框架,每个边缘节点配备一个调度智能体,负责将任务路由至本地ASP或邻近边缘节点。由于调度智能体仅基于局部观测决策,而任务卸载会耦合各智能体的资源状态与长期收益,因此将调度过程建模为合作式去中心化部分可观测马尔可夫决策过程(Dec-POMDP)。为求解该问题,我们提出了注意力引导的多智能体深度强化学习算法(A-MADiff),采用集中训练、去中心化执行范式。A-MADiff使用基于扩散的去中心化执行器生成多模态的可行调度偏好,并通过注意力引导的集中式评价器,从跨智能体状态中估计各智能体的价值,应对显存异构性。数值结果表明,A-MADiff在累积奖励上显著优于当前最优基线。
原文摘要 · Abstract (English)
Artificial Intelligence-Generated Content (AIGC) services employ Generative AI (GenAI) models to automatically generate diverse content. Mobile AIGC networks host GenAI models on edge-located AIGC Service Providers (ASPs) to deliver low-latency and personalized AIGC services for mobile users. However, AIGC inference tasks typically occupy GPU memory until task completion, causing GPU memory exhaustion at serving ASPs and triggering out-of-memory failures rather than merely increasing service latency. Existing studies on AIGC task orchestration have largely overlooked GPU memory feasibility constraints. To address this issue, we develop a cooperative multi-agent orchestration framework, in which each edge node is equipped with a scheduling agent to route tasks to local ASPs or neighboring edge nodes. Since scheduling agents make decisions based only on local observations, while peer offloading couples their resource states and long-term utilities, we formulate the orchestration process as a cooperative Decentralized Partially Observable Markov Decision Process (Dec-POMDP). To solve the Dec-POMDP, we propose an \underline{A}ttention-guided \underline{M}ulti-\underline{A}gent deep reinforcement learning algorithm with \underline{Diff}usion policies (A-MADiff) under the centralized training with a decentralized execution paradigm. A-MADiff employs diffusion-based decentralized actors to generate multi-modal preferences over feasible orchestration actions, and an attention-guided centralized critic to estimate per-agent values from cross-agent states under GPU memory heterogeneity. Numerical results demonstrate that A-MADiff significantly improves the cumulative reward over the state-of-the-art baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。