用预训练信念表示提升数据采集效率,通用且省样本。
Efficient Adaptive Data Acquisition via Pretrained Belief Representations

- 用预训练模型提取信念状态,分离表示与策略学习。
- 在多个任务中表现优于现有方法,训练样本减少70%以上。
- 适合需要高效采样的贝叶斯实验设计、优化与主动学习场景。
自适应数据采集策略的学习仍具挑战:基于后验的方法依赖代理模型和近似后验,易出现偏差;直接策略学习方法则仅依赖历史观测,无法利用已有模型表示,导致学习困难。我们提出基于信念表示的策略学习(POLAR),核心思想是最优数据采集仅依赖于历史观测所构成的充分信念状态。POLAR通过将预训练的预测基础模型作为信念状态编码器,将策略头训练在其表示之上,实现表示学习与策略学习的解耦。该方法构建了一个统一的摊销化策略学习框架,适用于贝叶斯实验设计、贝叶斯优化和主动学习,仅在任务特定效用函数上有所差异。实验证明,POLAR在多种任务中均超越现有先进摊销方法,且所需训练样本显著减少,显著提升了摊销数据采集的可扩展性与效率。
原文摘要 · Abstract (English)
Learning effective policies for adaptive data acquisition remains challenging: posterior-based methods rely on surrogate models and posterior approximations that can be misspecified or biased, while direct policy-learning methods map from historical observations and fail to exploit available model representations, making learning harder. We introduce policy learning with belief representations (POLAR), based on the insight that optimal data acquisition depends on the observation history only through a sufficient belief state. Specifically, POLAR decouples representation learning from policy learning by leveraging pretrained predictive foundation models as belief-state encoders, training a policy head on top of their representations. This yields a simple, unified amortised policy learning framework for Bayesian experimental design, Bayesian optimisation, and active learning, differing only in the task-specific utility used to train the policy. Empirically, we find that POLAR outperforms state-of-the-art amortised methods across diverse tasks while requiring far fewer training samples, demonstrating a significant step in the scalability and efficiency of amortised data acquisition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。