用数字孪生建模企业AI代理,提升决策能力
A Context Engineering Framework for Improving Enterprise AI Agents based on Digital-Twin MDP
- 构建数字孪生马尔可夫决策过程,抽象代理推理行为
- 通过对比逆强化学习从低质量数据中学习可靠奖励函数
- 适合需要高可靠性决策的企业AI系统优化
尽管企业在自动化与决策中广泛应用大模型代理,但其实际部署和性能提升仍受限于数据质量与数量不足、复杂现实推理需求、自对弈困难以及缺乏可靠反馈信号。为此,我们提出一种轻量级、模型无关的框架,通过离线强化学习改进基于大语言模型的企业代理。提出的基于数字孪生马尔可夫决策过程(DT-MDP-CE)框架包含三部分:(1) 数字孪生马尔可夫决策过程(DT-MDP),将代理推理行为抽象为有限马尔可夫决策过程;(2) 基于DT-MDP的鲁棒对比逆强化学习,高效估计可靠奖励函数,并从混合质量的离线轨迹中导出策略;(3) 强化学习引导的上下文工程,利用整合后得到的策略优化代理决策行为。以企业领域中的IT自动化任务为例,实验表明该框架在多种评估设置下均显著优于基线代理,具备良好的泛化能力。
原文摘要 · Abstract (English)
Despite rapid progress in AI agents for enterprise automation and decision-making, their real-world deployment and further performance gains remain constrained by limited data quality and quantity, complex real-world reasoning demands, difficulties with self-play, and the lack of reliable feedback signals. To address these challenges, we propose a lightweight, model-agnostic framework for improving LLM-based enterprise agents via offline reinforcement learning (RL). The proposed Context Engineering via DT-MDP (DT-MDP-CE) framework comprises three key components: (1) A Digital-Twin Markov Decision Process (DT-MDP), which abstracts the agent's reasoning behavior as a finite MDP; (2) A robust contrastive inverse RL, which, armed with the DT-MDP, to efficiently estimate a well-founded reward function and induces policies from mixed-quality offline trajectories; and (3) RL-guided context engineering, which uses the policy obtained from the integrated process of (1) and (2), to improve the agent's decision-making behavior. As a case study, we apply the framework to a representative task in the enterprise-oriented domain of IT automation. Extensive experimental results demonstrate consistent and significant improvements over baseline agents across a wide range of evaluation settings, suggesting that the framework can generalize to other agents sharing similar characteristics in enterprise environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。