物理AI agent通过主动推断实现测试时持续进化,应对未知场景。
Active Inference as the Test-Time Scaling Law for Physical AI Agents
- 基于主动推断原理,动态更新策略以应对未见环境
- 在自动驾驶仿真中提升泛化能力,推理效率提高36%以上
- 突破模型规模限制,随真实世界经验持续学习
本文提出一种新型物理人工智能(AI)代理的测试时扩展定律。该定律使物理AI代理能够利用其世界模型在测试时进行推理,从而在未预见的情境下实现泛化。该扩展定律基于主动推断的第一性原理,赋予代理在真实世界中生存的通用目标,其具体任务目标由此被包含其中。主动推断通过提供解决预测误差的推理机制,在代理遭遇训练分布外的意外情况时实现非平稳环境下的泛化。所提出的扩展定律通过在测试时动态更新代理策略来捕捉这一过程,该策略更新建模为软贝叶斯推理,信念依据降低预期预测误差的可允许策略作为似然进行更新。所得后验策略具有生物学解释,重现了大脑基底节与前额叶皮层在测试时的激活机制。为求解此解析不可行问题,开发了最小化自由能界变分推断方法,该方法扩展至支持训练外学习,通过在测试时解决的新实例强化策略与世界模型。与受限于模型规模和训练数据的现有扩展定律不同,该方法随物理智能体的连续真实世界经验而扩展。自动驾驶任务的仿真结果表明,该方案优于无模型Q学习和基于贝叶斯的模型强化学习,在未预见情境中表现出更强泛化能力,并将推理效率提升超过36%。
原文摘要 · Abstract (English)
In this paper, a novel test-time scaling law for physical artificial intelligence (AI) agents is introduced. This scaling law enables physical AI agents to reason with their world models to generalize in unforeseen scenarios at test time. The derived scaling law is grounded in the first principle of active inference, which equips agents with the general objective to survive in the real world, under which their specific task objectives are subsumed. Active inference achieves this by providing the reasoning to resolve prediction errors that arise when the agent encounters unforeseen situations outside its training distribution, enabling generalization in non-stationary environments. The proposed scaling law captures this by dynamically updating the agent's policy with this reasoning at test time. This policy update is modeled as a soft Bayesian inference process in which beliefs about the policy are updated using the reasoning that reduces expected prediction errors under allowable policies as a likelihood. The resulting posterior policy admits a biological interpretation, recovering the scaling mechanism that engages the brain's basal ganglia and prefrontal cortex at test time. To solve this analytically intractable problem, a variational inference solution minimizing free energy bounds is developed. This solution extends to enable learning beyond training by reinforcing new instances, resolved at test time, in both the policy and world model. Unlike existing scaling laws constrained by model size and training data, the derived solution scales with the continuous real-world experience of a physical AI agent. Simulation results on an autonomous driving task demonstrate that the proposed solution outperforms model-free Q-learning and model-based Bayesian reinforcement learning, achieving robust generalization to unforeseen scenarios while improving inference efficiency by over 36%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。