arXiv:2602.02760cs.CLcs.LG2026-02

测试大模型智能体在真实复杂环境下的适应能力,发现其表现远低于理论水平。

From Task Solving to Robust Real-World Adaptation in LLM Agents

  • 设计四类真实场景测试智能体的持续适应能力
  • 模型在长时序、高不确定性下性能显著下降,排名不固定
  • 无需明确指令也能自发权衡目标、效率与风险,具隐性目标推理能力

大型语言模型正作为可规划、调用工具并执行长期任务的智能体部署。然而现有评估常假设理想接口:规则明确稳定、工具可靠、目标清晰——这高估了实际应用准备度。现实中,智能体面临规则模糊、信号噪声、环境动态变化及多利益相关方的隐性目标。挑战不仅是完成任务,更是在执行中持续适应:判断可信度、识别需求、决定验证时机、选择退让或升级策略。本文在基于网格的游戏环境中,针对部分可观测、动态环境、噪声信号和状态漂移四种操作场景,对五种先进大模型智能体进行压力测试。任务虽简单但需长期执行,违反理想接口假设但仍可解,迫使智能体推断规则、付费获取信息、适应内外部变化,并在噪声中谨慎行动。结果表明,各模型在任务解决与部署鲁棒性间存在巨大差距;随着网格尺寸和执行时长增加,性能普遍下降,但排名不稳定:弱模型在适配不确定性的场景中可能超越强模型。尽管无显式指令,智能体仍自发权衡完成率、效率与惩罚规避,暗示具备部分目标推理能力。消融实验与特征分析揭示了模型特有的敏感性和失效原因,推动未来研究聚焦于验证机制、安全动作选择与部分可观测条件下的目标推断。

原文摘要 · Abstract (English)

Large language models are increasingly deployed as specialized agents that plan, call tools, and take actions over extended horizons. Yet many existing evaluations assume a "clean interface" where dynamics are specified and stable, tools and sensors are reliable, and success is captured by a single explicit objective-often overestimating real-world readiness. In practice, agents face underspecified rules, unreliable signals, shifting environments, and implicit, multi-stakeholder goals. The challenge is therefore not just solving tasks, but adapting while solving: deciding what to trust, what is wanted, when to verify, and when to fall back or escalate. We stress-test deployment-relevant robustness under four operational circumstances: partial observability, dynamic environments, noisy signals, and dynamic agent state. We benchmark agentic LLMs in a grid-based game with a simple goal but long-horizon execution. Episodes violate clean-interface assumptions yet remain solvable, forcing agents to infer rules, pay for information, adapt to environmental and internal shifts, and act cautiously under noise. Across five state-of-the-art LLM agents, we find large gaps between nominal task-solving and deployment-like robustness. Performance generally degrades as grid size and horizon increase, but rankings are unstable: weaker models can beat stronger ones when strategy matches the uncertainty regime. Despite no explicit instruction, agents trade off completion, efficiency, and penalty avoidance, suggesting partial objective inference. Ablations and feature analyses reveal model-specific sensitivities and failure drivers, motivating work on verification, safe action selection, and objective inference under partial observability, noise, and non-stationarity.

智能体鲁棒性大模型适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。