arXiv:2505.19867cs.LGcs.AI2025-05被引 2

用生成模型实现长时延环境下的端到端决策,无需人工奖励或耗时规划。

Deep Active Inference Agents for Delayed and Long-Horizon Environments

  • 用多步隐状态转移一次性预测整段未来
  • 单步梯度更新实现百步级长时程规划
  • 适合工业级复杂延迟任务,无需人工设计奖励

随着世界模型智能体的成功,基于可微分模型的模型化强化学习在多样任务中实现了高效控制。主动推理(AIF)提供了一种神经科学基础的互补范式,将感知、学习与行动统一于一个由生成模型驱动的概率框架中。然而,现有AIF智能体仍依赖精确的即时预测和穷尽式规划,在需长时间跨度(数十至数百步)的延迟环境中表现受限。此外,多数现有方法仅在机器人或视觉基准上评估,难以反映真实工业场景的复杂性。本文提出一种生成策略架构:(i) 多步隐状态转移使生成模型能一次性预测整个时间窗口;(ii) 集成策略网络接收期望自由能梯度;(iii) 采用交替优化方案从回放缓冲区更新模型与策略;(iv) 单步梯度实现长时程规划,消除控制循环中的穷尽式规划。我们在模拟真实工业场景的延迟与长时程环境中评估该智能体,结果表明,所提方法有效结合世界模型与AIF形式,实现无需人工奖励或昂贵规划的端到端概率控制器,在延迟、长时程环境下具备高效决策能力。

原文摘要 · Abstract (English)

With the recent success of world-model agents, which extend the core idea of model-based reinforcement learning by learning a differentiable model for sample-efficient control across diverse tasks, active inference (AIF) offers a complementary, neuroscience-grounded paradigm that unifies perception, learning, and action within a single probabilistic framework powered by a generative model. Despite this promise, practical AIF agents still rely on accurate immediate predictions and exhaustive planning, a limitation that is exacerbated in delayed environments requiring plans over long horizons, tens to hundreds of steps. Moreover, most existing agents are evaluated on robotic or vision benchmarks which, while natural for biological agents, fall short of real-world industrial complexity. We address these limitations with a generative-policy architecture featuring (i) a multi-step latent transition that lets the generative model predict an entire horizon in a single look-ahead, (ii) an integrated policy network that enables the transition and receives gradients of the expected free energy, (iii) an alternating optimization scheme that updates model and policy from a replay buffer, and (iv) a single gradient step that plans over long horizons, eliminating exhaustive planning from the control loop. We evaluate our agent in an environment that mimics a realistic industrial scenario with delayed and long-horizon settings. The empirical results confirm the effectiveness of the proposed approach, demonstrating the coupled world-model with the AIF formalism yields an end-to-end probabilistic controller capable of effective decision making in delayed, long-horizon settings without handcrafted rewards or expensive planning.

主动推理长时程规划生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。