arXiv:2602.07215eess.SYcs.AI2026-02被引 1

用多智能体协作优化移动端边缘网络中的生成式AI推理,兼顾速度与公平性。

Multi-Agentic AI for Fairness-Aware and Accelerated Multi-modal Large Model Inference in Real-world Mobile Edge Networks

  • 设计多智能体系统,通过自然语言推理动态调度提示和模型部署。
  • 平均延迟降低80%以上,公平性指标达0.90,优于现有基线。
  • 无需微调即可快速适应新场景,适合实际边缘部署的通用方案。

生成式人工智能(GenAI)在自然语言处理和内容创作中已带来变革,但集中式推理仍受限于高延迟、定制化能力弱及隐私问题。将大模型(LMs)部署于移动边缘网络成为有前景的解决方案。然而,这也带来了新挑战:异构多模态大模型资源需求和推理速度差异大,提示与输出模态多样导致编排复杂,且边缘基础设施资源有限,难以支持并发执行。为此,我们提出一种面向延迟与公平性的多智能体AI框架,用于移动边缘网络中的多模态大模型推理。该框架包含长期规划智能体、短期提示调度智能体以及多个节点级模型部署智能体,均由基础语言模型驱动。这些智能体通过运行时遥测数据与历史经验进行自然语言推理,协同优化提示路由与模型部署。为评估性能,我们构建了一个城市级测试平台,支持网络监控、容器化模型部署、服务器内资源管理与跨服务器通信。实验表明,该方案平均延迟降低超80%,公平性(归一化Jain指数)提升至0.90,优于其他基线。此外,该方案无需微调即可快速适应,为边缘环境中生成式AI服务优化提供可泛化的解决方案。

原文摘要 · Abstract (English)

Generative AI (GenAI) has transformed applications in natural language processing and content creation, yet centralized inference remains hindered by high latency, limited customizability, and privacy concerns. Deploying large models (LMs) in mobile edge networks emerges as a promising solution. However, it also poses new challenges, including heterogeneous multi-modal LMs with diverse resource demands and inference speeds, varied prompt/output modalities that complicate orchestration, and resource-limited infrastructure ill-suited for concurrent LM execution. In response, we propose a Multi-Agentic AI framework for latency- and fairness-aware multi-modal LM inference in mobile edge networks. Our solution includes a long-term planning agent, a short-term prompt scheduling agent, and multiple on-node LM deployment agents, all powered by foundation language models. These agents cooperatively optimize prompt routing and LM deployment through natural language reasoning over runtime telemetry and historical experience. To evaluate its performance, we further develop a city-wide testbed that supports network monitoring, containerized LM deployment, intra-server resource management, and inter-server communications. Experiments demonstrate that our solution reduces average latency by over 80% and improves fairness (Normalized Jain index) to 0.90 compared to other baselines. Moreover, our solution adapts quickly without fine-tuning, offering a generalizable solution for optimizing GenAI services in edge environments.

多智能体边缘计算生成式AI公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。