arXiv:2602.17976cs.LGcs.AI2026-02中稿 · ICML

提出连续空间纯探索新方法,高效找近优解

In-Context Pure Exploration in Continuous Decision Spaces

  • 基于贝叶斯框架设计可元训练的探索模型
  • 在连续动作空间中实现ε-最优推荐,仅需少量查询
  • 适合连续决策场景,如函数最小值估计或区域定位

在主动序列测试(又称纯探索)中,学习者需自适应地获取信息,以最少的查询识别未知的真实假设。该问题在多臂赌博机的最佳臂识别及广义搜索问题中有重要应用。然而,在许多现代场景中,假设空间是连续的:例如在连续动作赌博机中寻找近似最优动作、定位包含于目标区域的ε-球,或从噪声观测中估计函数最小值。现有方法多为频率学派且依赖特定模型,而基于学习的方法此前仅限于有限推荐空间。本文提出C-ICPE,一种理论引导的贝叶斯固定置信度纯探索模型,支持连续推荐。C-ICPE在任务先验上元训练序列架构,联合学习探索、停止与推荐策略。推理时,它无需参数更新即可主动收集证据并识别ε-最优推荐。

原文摘要 · Abstract (English)

In active sequential testing, also termed pure exploration, a learner is tasked with the goal to adaptively acquire information so as to identify an unknown ground-truth hypothesis with as few queries as possible. This problem has several motivating applications, including Best-Arm Identification (BAI) in bandits, where actions index hypotheses, and generalized search problems, where strategically chosen queries reveal partial information about a hidden label. In many modern settings, however, the hypothesis, or recommendation space, is continuous: for example, identifying a near optimal action in a continuous-armed bandit, localizing an $ε$-ball contained in a target region, or estimating the minimizer of a function from noisy observations. Existing methods are predominantly frequentist and model-specific, while learned approaches have been limited to finite recommendation spaces. We introduce C-ICPE, a theory-guided learned model for Bayesian fixed-confidence pure exploration with continuous recommendations. C-ICPE meta-trains sequential architectures over a task prior to jointly learn exploration, stopping and recommendations strategies. At inference time, it actively gathers evidence on tasks and identifies an $ε$-optimal recommendation without parameter updates.

纯探索连续空间贝叶斯优化元学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。