arXiv:2508.05619cs.AInlin.AO2025-08被引 2

用内在动机替代人工奖励,让AI自主学习并适应环境变化。

The Missing Reward: Active Inference in the Era of Experience

  • 以自由能最小化代替外部奖励,实现探索与利用的统一平衡
  • 结合大语言模型作为世界模型,使智能体从自生成数据中高效学习
  • 适合追求自主性、低依赖人工干预的下一代AI系统研究者

本文指出,主动推理(Active Inference, AIF)为构建能够从经验中自主学习而无需持续人工奖励设计的智能体提供了关键基础。当前人工智能在高质量数据耗尽、依赖大规模人力进行奖励设计的背景下,面临显著的可扩展性挑战,阻碍了真正自主智能的发展。‘经验时代’的愿景虽具前景,但仍严重依赖人工设计的奖励函数,实质上只是将瓶颈从数据整理转移到奖励设计。这暴露了我们所称的‘具身代理鸿沟’:现有AI无法自主形成、调整并追求目标以应对环境变化。我们提出AIF可通过用内在的自由能最小化驱动力取代外部奖励信号,使智能体自然实现探索与利用的平衡。通过将大语言模型作为生成式世界模型,与AIF的严谨决策框架结合,可构建出既能高效从经验中学习,又能保持与人类价值观对齐的智能体。这一融合为具有自主性的智能系统提供了一条兼具计算与物理约束的可行路径。

原文摘要 · Abstract (English)

This paper argues that Active Inference (AIF) provides a crucial foundation for developing autonomous AI agents capable of learning from experience without continuous human reward engineering. As AI systems begin to exhaust high-quality training data and rely on increasingly large human workforces for reward design, the current paradigm faces significant scalability challenges that could impede progress toward genuinely autonomous intelligence. The proposal for an ``Era of Experience,'' where agents learn from self-generated data, is a promising step forward. However, this vision still depends on extensive human engineering of reward functions, effectively shifting the bottleneck from data curation to reward curation. This highlights what we identify as the \textbf{grounded-agency gap}: the inability of contemporary AI systems to autonomously formulate, adapt, and pursue objectives in response to changing circumstances. We propose that AIF can bridge this gap by replacing external reward signals with an intrinsic drive to minimize free energy, allowing agents to naturally balance exploration and exploitation through a unified Bayesian objective. By integrating Large Language Models as generative world models with AIF's principled decision-making framework, we can create agents that learn efficiently from experience while remaining aligned with human values. This synthesis offers a compelling path toward AI systems that can develop autonomously while adhering to both computational and physical constraints.

主动推理自主智能大模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。