动态调整边缘生成推理资源,实时应对设备变化和负载波动。
$E^3$-Agent: An Executable and Evolving Agent for Resource Management of Edge Generative Inference

- 用快速路由+慢速大模型控制器分离架构,毫秒级决策并在线自适应。
- 相比静态基线,平均延迟降低65%-73%,接近理想情况的90%以上性能。
- 适合需要高稳定性和低延迟的边缘生成内容应用,如实时视频生成。
边缘生成式推理部署面临两大现实挑战:部署时各设备各模型的性能未知且持续变化,受用户语义事件、后台负载和设备更替影响。因此,离线调优的静态资源管理器易失效且维护成本高。本文提出E³-Agent,一种可执行且可演进的边缘AI生成内容(AIGC)资源管理代理。该代理将毫秒级调度的快速路径与基于事件的慢速大语言模型(LLM)元控制器分离,通过工具接口提供风险控制、路由配置和快速性能校准等显式控制面,以应对模式漂移。代理在执行中在线学习反馈,持续适应未知且随时间变化的服务时间映射。我们在基于MLPerf设备-模型测量先验的离散事件模拟器中评估,覆盖冷启动预热及三种动态场景:语义动态、设备更替与隐性漂移。在各类动态场景下,E³-Agent相较最优静态基线平均延迟降低65%-73%,性能保持在用于评估的在线全信息Oracle的7%-10%以内,并有效抑制了语义退化下的卡顿率。
原文摘要 · Abstract (English)
Edge deployments of generative inference increasingly face two practical realities: per-device per-model performance is often unknown at deployment time, and it is non-stationary due to user-driven semantic events, background load, and device churn. Consequently, a resource manager that is tuned offline under a fixed regime can become brittle and expensive to maintain. This paper presents $E^3$-Agent, an executable and evolving agent for edge artificial intelligence generated content (AIGC) resource management. $E^3$-Agent separates a fast-path router that makes millisecond-level dispatch decisions from a slow-path, event-driven large language model (LLM) meta-controller that mitigates regime shifts through a small, explicit control surface exposed via a tool interface, including risk gating, router configuration, and rapid performance calibration. The agent learns online from execution feedback and continuously adapts to unknown and time-varying service-time mappings. We evaluate $E^3$-Agent in a discrete-event simulator driven by MLPerf-derived device-model measurement priors, covering cold-start warmup and three dynamic regimes: semantic dynamics, device churn, and hidden drift. Across the dynamic scenarios, $E^3$-Agent reduces average latency by 65%-73% compared to the best static baseline, stays within 7%-10% of an online full-information Oracle used for evaluation, and effectively suppresses stutter rate under semantic degradation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。