arXiv:2512.04745math.OCcs.AI2025-12被引 2

用自由能最小化实现智能体策略的自动组合,让行为更灵活可解释。

Neural Policy Composition from Free Energy Minimization

  • 从自由能最小化推导出策略组合的连续时间梯度流
  • 在多智能体、决策任务等场景中稳定收敛并匹配最优组合
  • 适合研究智能行为机制与神经动力学建模的研究者

灵活组合已有技能以执行智能行为是自然智能的特征。这种组合灵活性通常归因于上下文相关的门控机制,决定多个策略或行为基元如何结合。然而,尽管有大量研究,这些门控规则应遵循的规范性目标以及实现它们的神经计算机制仍不明确。现有方法通常依赖预设的门控设计,且受限于特定架构、学习范式或数据集。本文提出一个规范性框架,使策略组合源于变分自由能最小化,为门控提供普适而严谨的目标。基于此框架,我们推导出一种连续时间梯度流,其轨迹保证收敛至基元最优组合,并具有明确收敛速率。进一步表明,该动态可被机制化地实现为具有上下文敏感局部交互的软竞争递归电路。我们在多智能体系统的涌现蜂群行为、人类在老虎机任务中的决策、分层架构的控制基准测试中评估模型。在各类场景中,模型提供了可解释的策略组合机制,再现关键行为特征,揭示数据内在规律,表现达到或优于现有模型。

原文摘要 · Abstract (English)

The ability to flexibly compose previously acquired skills to execute intelligent behaviors is a hallmark of natural intelligence. Such compositional flexibility is often attributed to context-dependent gating mechanisms that determine how multiple policies or behavioral primitives are combined. Yet, despite remarkable efforts, the normative objective from which such gating rules should arise, and the neural computations capable of implementing them, remain unclear. Existing approaches typically rely on prespecified design choices for the gating rules, and remain tied to specific architectures, learning paradigms, or datasets. Here, we introduce a normative framework in which policy composition emerges from the minimization of a variational free energy, providing a principled and broadly applicable objective for gating. Based on this framework, we derive a continuous-time gradient flow whose trajectories are guaranteed to converge, with explicit rate, to the optimal composition of primitives. We further show that this dynamics admits a mechanistic neural implementation as a soft-competitive recurrent circuit with context-sensitive local interactions. We evaluate the model on emerging flocking behaviors in multi-agent systems, human decision-making in bandit tasks, and control benchmarks in layered architectures. Across these settings, the model provides interpretable mechanistic accounts of policy composition, reproduces key behavioral signatures, yields insights into data, and matches or outperforms established models.

强化学习策略组合神经动力学自由能原理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。