arXiv:2505.12247cs.NIcs.AI2025-05被引 22

用大模型提升网络代理对用户意图的理解,优化服务体验。

LAMeTA: Intent-Aware Agentic Network Optimization via a Large AI Model-Empowered Two-Stage Approach

论文配图:LAMeTA: Intent-Aware Agentic Network Optimization via a Large AI Model-Empowered Two-Stage Approach
图 1 · 摘自论文原文
  • 通过意图引导的知识蒸馏,将大模型能力迁移到轻量边缘模型。
  • 结合自然语言意图生成偏好向量,使强化学习更精准优化服务质量。
  • 适合研究智能网络优化或大模型应用的开发者与工程师。

当前生成式AI正重塑多个领域,使机器能够跨模态生成内容。随着生成式AI演变为具备推理、协作和交互能力的自主代理,它们越来越多地部署在网络基础设施上以自动服务人类。这一新兴范式称为代理网络,因需融入用户以自然语言表达的主观意图而带来新的优化挑战。传统通用深度强化学习(DRL)难以捕捉意图语义并动态调整策略,导致性能不佳。本文提出LAMeTA,一种基于大模型(LAM)的两阶段意图感知代理网络优化方法。首先,提出面向意图的知识蒸馏(IoKD),高效将资源密集型大模型的意图理解能力迁移至轻量级边缘大模型(E-LAM),以服务终端用户。其次,构建共生强化学习(SRL),将E-LAM与基于策略的DRL框架结合。在SRL中,E-LAM将自然语言用户意图转化为结构化偏好向量,指导状态表示与奖励设计;DRL则根据实时网络条件优化生成服务功能链组合与E-LAM选择,从而提升主观体验质量(QoE)。在包含81个代理的代理网络中进行的大量实验表明,IoKD使意图预测均方误差降低最高达22.5%,而SRL在最大化意图感知QoE方面比传统通用DRL最高提升23.5%。

原文摘要 · Abstract (English)

Nowadays, Generative AI (GenAI) reshapes numerous domains by enabling machines to create content across modalities. As GenAI evolves into autonomous agents capable of reasoning, collaboration, and interaction, they are increasingly deployed on network infrastructures to serve humans automatically. This emerging paradigm, known as the agentic network, presents new optimization challenges due to the demand to incorporate subjective intents of human users expressed in natural language. Traditional generic Deep Reinforcement Learning (DRL) struggles to capture intent semantics and adjust policies dynamically, thus leading to suboptimality. In this paper, we present LAMeTA, a Large AI Model (LAM)-empowered Two-stage Approach for intent-aware agentic network optimization. First, we propose Intent-oriented Knowledge Distillation (IoKD), which efficiently distills intent-understanding capabilities from resource-intensive LAMs to lightweight edge LAMs (E-LAMs) to serve end users. Second, we develop Symbiotic Reinforcement Learning (SRL), integrating E-LAMs with a policy-based DRL framework. In SRL, E-LAMs translate natural language user intents into structured preference vectors that guide both state representation and reward design. The DRL, in turn, optimizes the generative service function chain composition and E-LAM selection based on real-time network conditions, thus optimizing the subjective Quality-of-Experience (QoE). Extensive experiments conducted in an agentic network with 81 agents demonstrate that IoKD reduces mean squared error in intent prediction by up to 22.5%, while SRL outperforms conventional generic DRL by up to 23.5% in maximizing intent-aware QoE.

网络优化大模型强化学习意图理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。