arXiv:2603.06669cs.NIcs.AI2026-03

用自模仿强化学习统一优化边缘AI微服务的部署与路由。

Hybrid Orchestration of Edge AI and Microservices via Graph-based Self-Imitation Learning

  • 基于图注意力网络建模服务拓扑,结合自模仿学习提升策略探索效率。
  • 在真实负载下将端到端延迟降低32%以上,资源利用率提升28%。
  • 适合需要低延迟、高并发的边缘AI系统开发者与架构师参考。

现代边缘AI应用越来越多地依赖微服务架构,将AI服务与传统微服务整合为复杂请求链,并面临严格的延迟要求。有效编排这些异构服务对保障低延迟性能至关重要,但因资源需求多样且在资源受限的边缘环境中存在强操作耦合性而极具挑战。尤其服务间频繁交互紧密耦合部署与路由决策,现有方法却孤立优化二者,导致系统性能根本不足。本文提出SIL-GPO,一种强化学习框架,用于优化边缘AI微服务系统的混合编排。SIL-GPO将编排问题建模为序列决策任务,利用图注意力网络编码服务拓扑与路由依赖关系至智能体状态表示。同时,将自模仿学习引入近端策略优化,使智能体优先复用高奖励轨迹,引导策略更新向全局最优解逼近,克服稀疏奖励与大规模组合动作空间下的探索困境。我们在基于实际追踪数据的边缘AI工作负载上进行大量实验,结果表明,SIL-GPO显著降低端到端服务延迟,并优于当前最先进的启发式、元启发式及深度强化学习基线。本框架为边缘环境下AI服务与微服务的高效编排提供了统一且可扩展的解决方案,推动低延迟、高性能边缘AI部署的发展。

原文摘要 · Abstract (English)

Modern edge AI applications increasingly rely on microservice architectures that integrate both AI services and conventional microservices into complex request chains with stringent latency requirements. Effectively orchestrating these heterogeneous services is crucial for ensuring low-latency performance, yet remains challenging due to their diverse resource demands and strong operational interdependencies under resource-constrained edge environments. In particular, frequent interactions between services tightly couple deployment and routing decisions, yet existing approaches optimize them in isolation, leading to fundamentally inadequate system performance.In this paper, we propose SIL-GPO, a reinforcement learning framework that optimizes hybrid orchestration for edge AI microservice systems. SIL-GPO formulates the orchestration problem as a sequential decision-making task and leverages graph attention networks to encode service topologies and routing dependencies within the agent state representation. Moreover, SIL-GPO integrates a self-imitation learning strategy into proximal policy optimization, enabling the agent to prioritize and reuse high-reward trajectories. This guides policy updates towards globally promising solutions that standard RL often fails to discover under sparse rewards and large combinatorial action spaces. We conduct extensive experiments on trace-driven edge AI workloads, demonstrating that SIL-GPO significantly reduces end-to-end service latency and enhances resource utilization compared to state-of-the-art heuristic, metaheuristic, and deep RL baselines. Our framework offers a unified and scalable solution for efficient orchestration of AI services and microservices in the edge, paving the way for low-latency, high-performance edge AI deployments.

边缘计算强化学习微服务编排

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。