用可复用的价值配置实现灵活策略切换,让智能体在变化环境中高效决策。
Active Inference with Reusable State-Dependent Value Profiles
- 用隐藏状态关联的参数包动态生成控制策略,避免为每种情境单独设置参数
- 模型在逆转学习任务中表现最优,信息准则相差约100点,支持结构可识别性
- 发现策略调整主要依赖政策先验而非精度调节,体现基于信念的渐进式控制
在多变环境中,智能体需根据隐含状态在不同价值控制模式间切换,但为每种情境独立维护偏好、策略偏置和动作置信参数不可行。本文提出价值配置:一组可复用的小规模参数包(结果偏好、策略先验、策略精度),分配给生成模型中的隐状态。随着后验信念逐次更新,有效控制参数通过信念加权混合自然生成,实现无需独立参数的条件策略调用。在概率逆转学习任务中,通过交叉验证对数似然与信息准则对比静态精度、熵耦合动态精度及配置基模型,结果显示配置基模型显著更优(约100点AIC差异)。参数恢复分析表明即使在噪声观测下也能识别结构。模型推断进一步揭示:该任务中的自适应控制主要由策略先验调节驱动,而非策略精度;信念依赖的渐进式配置过程符合状态条件控制,而非单纯由不确定性驱动。总体而言,可复用的价值配置为多变环境中的信念条件价值控制提供了可行计算机制,并产生可检验的信念依赖控制与行为灵活性信号。
原文摘要 · Abstract (English)
Adaptive behavior in volatile environments requires agents to switch among value-control regimes across latent contexts, but maintaining separate preferences, policy biases, and action-confidence parameters for every situation is intractable. We introduce value profiles: a small set of reusable bundles of value-related parameters (outcome preferences, policy priors, and policy precision) assigned to hidden states in a generative model. As posterior beliefs over states evolve trial by trial, effective control parameters arise via belief-weighted mixing, enabling state-conditional strategy recruitment without requiring independent parameters for each context. We evaluate this framework in probabilistic reversal learning, comparing static-precision, entropy-coupled dynamic-precision, and profile-based models using cross-validated log-likelihood and information criteria. Model comparison favors the profile-based model over simpler alternatives (about 100-point AIC differences), and parameter-recovery analyses support structural identifiability even when context must be inferred from noisy observations. Model-based inference further suggests that adaptive control in this task is driven primarily by modulation of policy priors rather than policy precision, with gradual belief-dependent profile recruitment consistent with state-conditional (not purely uncertainty-driven) control. Overall, reusable value profiles provide a tractable computational account of belief-conditioned value control in volatile environments and yield testable signatures of belief-dependent control and behavioral flexibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。