arXiv:2602.06104cs.LGstat.ML2026-02

用主动推理统一学习与优化,让决策更智能且少试错。

Pragmatic Curiosity: A Unified Framework for Hybrid Learning and Optimization via Active Inference

  • 通过信息增益与损失预期权衡,动态选择最优查询点。
  • 在三种复杂场景中均降低决策风险,提升关键区域覆盖度。
  • 无需人为设定阶段规则,自动联合学习预测与偏好结构。

许多工程与科学流程依赖昂贵的黑箱评估,需在序列决策中同时提升任务性能并减少不确定性。贝叶斯优化(BO)与贝叶斯实验设计(BED)分别处理目标导向优化与信息获取,但对学习与优化耦合的混合场景指导有限。本文提出实用好奇心(PraC),一种基于主动推理的统一框架,用于混合学习与优化。PraC通过权衡任务相关隐变量的信息增益与结果的预期后悔值来评估候选查询。该框架揭示三个可配置设计:应澄清哪个隐变量、如何将任务价值编码为后悔值、信息增益与实用价值的交换强度。我们在三个递增复杂度的场景中实例化PraC:固定全局符号与已知下游损失的决策导向烟羽监测;引入局部符号与动态覆盖目标的定向主动搜索;未知偏好下的分层后悔学习复合贝叶斯优化。在这些场景中,PraC显著降低下游决策风险,改善关键结果区域覆盖率,并在不依赖任务特定分阶段规则的情况下,联合学习预测与偏好结构。

原文摘要 · Abstract (English)

Many engineering and scientific workflows rely on expensive black-box evaluations, requiring sequential decisions that must both improve task performance and reduce uncertainty. Bayesian optimization (BO) and Bayesian experimental design (BED) provide powerful but largely separate treatments of goal-directed optimization and information-seeking experimentation, leaving limited guidance for hybrid settings in which learning and optimization are intrinsically coupled. We propose Pragmatic Curiosity (PraC), a unified framework for hybrid learning and optimization via active inference. PraC evaluates candidate queries by trading information gain about a task-relevant latent symbol against an expected regret-based potential over outcomes. This formulation exposes three operational design choices: which latent quantity should be clarified, how task value is encoded as regret, and how strongly information gain should be exchanged against pragmatic value. We instantiate PraC across three regimes of increasing complexity: decision-oriented plume monitoring with fixed global symbols and known downstream losses, targeted active search with induced local symbols and evolving coverage goals, and composite Bayesian optimization with hierarchical regret learning under unknown preferences. Across these regimes, PraC reduces downstream decision risk, improves coverage of critical outcome regions, and jointly learns predictive and preference structures without relying on task-specific staging rules.

主动推理贝叶斯优化决策学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。