arXiv:2601.03905cs.AIcs.CL2026-01ACL被引 15

现有智能体难以有效利用世界模型进行前瞻决策。

Current Agents Fail to Leverage World Model as Tool for Foresight

  • 测试发现仅不足1%情况下调用模拟,15%出现预测使用错误。
  • 引入模拟后性能反而下降最多达5%,表明推理不一致。
  • 核心瓶颈在于何时模拟、如何解读结果及整合预见信息。

基于视觉语言模型的智能体面临需预判未来状态的任务,而非依赖短时推理。生成式世界模型提供潜在解决方案:智能体可将其作为外部模拟器,在行动前预演结果。本文实证检验当前智能体是否能有效利用此类世界模型提升认知能力。在多样化的智能体与视觉问答任务中,部分智能体极少调用模拟(少于1%),频繁误用预测轨迹(约15%),且在可使用或强制启用模拟时表现出不一致甚至性能下降(最高达5%)。归因分析表明,主要瓶颈在于智能体判断何时模拟、如何解读预测结果以及将预见信息融入下游推理的能力。研究强调需建立校准、策略性的世界模型交互机制,以推动未来智能体系统更可靠的前瞻认知发展。

原文摘要 · Abstract (English)

Agents built on vision-language models increasingly face tasks that demand anticipating future states rather than relying on short-horizon reasoning. Generative world models offer a promising remedy: agents could use them as external simulators to foresee outcomes before acting. This paper empirically examines whether current agents can leverage such world models as tools to enhance their cognition. Across diverse agentic and visual question answering tasks, we observe that some agents rarely invoke simulation (fewer than 1%), frequently misuse predicted rollouts (approximately 15%), and often exhibit inconsistent or even degraded performance (up to 5%) when simulation is available or enforced. Attribution analysis further indicates that the primary bottleneck lies in the agents' capacity to decide when to simulate, how to interpret predicted outcomes, and how to integrate foresight into downstream reasoning. These findings underscore the need for mechanisms that foster calibrated, strategic interaction with world models, paving the way toward more reliable anticipatory cognition in future agent systems.

智能体世界模型前瞻推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。