arXiv:2606.07017cs.AIcs.CL2026-06KDD被引 2

把大模型智能体的仿真到现实差距,用马尔可夫决策过程统一分析。

The Sim-to-Real Gap of Foundation Model Agents: A Unified MDP Perspective

论文配图:The Sim-to-Real Gap of Foundation Model Agents: A Unified MDP Perspective
图 1 · 摘自论文原文
  • 从观测、动作、转移、奖励四要素重构仿真到现实的差距问题
  • 多语言工具调用案例显示:语义正确但因观测差异导致无效操作
  • 倡导引入领域随机化等经典方法,推动建立标准化评测基准

基础模型智能体在真实世界决策中应用日益广泛,但面临严重的仿真到现实差距。尽管机器人学与经典控制已有成熟框架应对此问题,基础模型领域却将其视为全新挑战。本文将基础模型智能体的评估与训练差距形式化为经典的仿真到现实问题,完全基于马尔可夫决策过程的四个要素:观测、动作、转移与奖励。我们提出一个全面的研究议程,将经典差异映射至基础模型领域,并倡导采用领域随机化等既有解决方案。通过多语言工具调用案例展示,即使语义意图正确,观测空间差异仍会导致操作无效。最终目标是实现范式转变,建立统一术语体系与标准化压力测试基准,培育更可信的下一代真实世界智能体。

原文摘要 · Abstract (English)

Foundation model agents are increasingly deployed for real-world decision-making, but suffer from the sim-to-real gap. While robotics and classical control have mature frameworks to address this gap, the foundation model community is treating agent robustness as an entirely novel phenomenon. Our paper proposes formalizing the foundation model agent evaluation and training gap as a classical sim-to-real problem structured entirely around the four elements of a Markov Decision Process, including Observation, Action, Transition, and Reward. In this paper, we set a comprehensive research agenda that translates classical discrepancies into the foundation model domain and advocates for adopting established solutions like domain randomization. We provide concrete examples, such as a multilingual tool calling to demonstrate how severe observation space gaps lead to operationally invalid actions despite correct semantic intent. Ultimately, this agenda aims to drive a paradigm shift, yielding a unified vocabulary and standardized stress test benchmarks to foster a new generation of highly trustworthy agents for reliable real-world applications.

大模型智能体仿真实战决策过程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。