arXiv:2602.06029cs.LG2026-02被引 2

提出好奇心调控理论,让智能体既学得对又决策不后悔。

Curiosity is Knowledge: Self-Consistent Learning and No-Regret Optimization with Active Inference

  • 用好奇心系数统一控制探索与决策,确保学习一致
  • 仅需足够好奇心,就能实现无遗憾优化和后验一致
  • 适用于需要边学边做的复杂任务,如机器人规划

主动推理(AIF)通过最小化期望自由能(EFE)统一了探索与利用,借助好奇心系数平衡认知价值(信息增益)与实用价值(任务表现)。然而,这种平衡何时能同时保证一致学习与高效决策仍不明确:好奇心不足会导致短视利用并阻碍不确定性消除,而过度好奇心则引发不必要的探索与后悔。本文首次为最小化EFE的智能体建立了理论保障,证明只要满足‘足够好奇心’这一单一条件,即可同时实现自洽学习(贝叶斯后验一致性)与无遗憾优化(累积后悔有界)。分析揭示该机制依赖于初始不确定性、可辨识性及目标对齐程度,从而将AIF与经典贝叶斯实验设计和贝叶斯优化统一于同一理论框架。进一步将理论转化为实际调参指南,用于混合学习-优化问题,并通过真实世界实验验证。

原文摘要 · Abstract (English)

Active inference (AIF) unifies exploration and exploitation by minimizing the Expected Free Energy (EFE), balancing epistemic value (information gain) and pragmatic value (task performance) through a curiosity coefficient. Yet it has been unclear when this balance yields both coherent learning and efficient decision-making: insufficient curiosity can drive myopic exploitation and prevent uncertainty resolution, while excessive curiosity can induce unnecessary exploration and regret. We establish the first theoretical guarantee for EFE-minimizing agents, showing that a single requirement--sufficient curiosity--simultaneously ensures self-consistent learning (Bayesian posterior consistency) and no-regret optimization (bounded cumulative regret). Our analysis characterizes how this mechanism depends on initial uncertainty, identifiability, and objective alignment, thereby connecting AIF to classical Bayesian experimental design and Bayesian optimization within one theoretical framework. We further translate these theories into practical design guidelines for tuning the epistemic-pragmatic trade-off in hybrid learning-optimization problems, validated through real-world experiments.

主动推理好奇心贝叶斯优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。