arXiv:2512.12548cs.AIcs.LG2025-12

用世界模型让AI学会像生物一样高效觅食

World Models Unlock Optimal Foraging Strategies in Reinforcement Learning Agents

  • 给AI装上预测环境的世界模型,让它能提前规划离开资源区
  • 模型类智能体的离群行为符合生态最优理论(MVT)
  • 适合研究可解释AI与生物决策机制的学者

觅食中的补丁选择涉及在资源丰富区域停留多久后转向潜在更优区域的策略性决策。边际价值定理(MVT)常用于描述这一过程,提供了一种行为优化模型。尽管该理论广泛应用于行为生态学预测,但揭示生物觅食者实现最优决策的计算机制仍不明确。本文表明,配备学习型世界模型的人工觅食者自然收敛至符合MVT的策略。通过基于模型的强化学习代理构建简洁的环境预测表示,我们发现前瞻能力而非单纯奖励最大化驱动了高效离群行为。相较于标准无模型强化学习代理,这些模型类代理展现出与多种生物体相似的决策模式,表明预测性世界模型可成为更具可解释性与生物学基础的AI决策框架。总体而言,我们的研究凸显了生态最优原则对推动可解释且自适应人工智能的重要性。

原文摘要 · Abstract (English)

Patch foraging involves the deliberate and planned process of determining the optimal time to depart from a resource-rich region and investigate potentially more beneficial alternatives. The Marginal Value Theorem (MVT) is frequently used to characterize this process, offering an optimality model for such foraging behaviors. Although this model has been widely used to make predictions in behavioral ecology, discovering the computational mechanisms that facilitate the emergence of optimal patch-foraging decisions in biological foragers remains under investigation. Here, we show that artificial foragers equipped with learned world models naturally converge to MVT-aligned strategies. Using a model-based reinforcement learning agent that acquires a parsimonious predictive representation of its environment, we demonstrate that anticipatory capabilities, rather than reward maximization alone, drive efficient patch-leaving behavior. Compared with standard model-free RL agents, these model-based agents exhibit decision patterns similar to many of their biological counterparts, suggesting that predictive world models can serve as a foundation for more explainable and biologically grounded decision-making in AI systems. Overall, our findings highlight the value of ecological optimality principles for advancing interpretable and adaptive AI.

强化学习世界模型觅食策略可解释AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。