arXiv:2501.01248cs.LG2025-01被引 1

用分布分歧提升回归主动学习,区分不确定性类型

Bayesian Active Learning By Distribution Disagreement

  • 基于归一化流模型,通过分布分歧设计新查询策略
  • 在4个数据集上实现当前最优性能,支持多种查询规模
  • 适合需要精准不确定性估计的回归场景

回归任务的主动学习因难以衡量不确定性而研究不足。由于归一化流能提供完整的预测分布而非点估计,可直接应用熵或置信度最低等传统主动学习启发式方法。然而我们发现,这些方法在基于归一化流的池采样主动学习中表现不佳,需更复杂的算法来区分偶然性与认知性不确定性。本文提出BALSA,是对BALD算法的适配改进,专用于归一化流的回归任务。本工作将归一化流的不确定性量化研究扩展至真实世界数据,并支持多种采集函数与查询规模。在4个不同数据集和2种架构上,BALSA均取得当前最优结果。

原文摘要 · Abstract (English)

Active Learning (AL) for regression has been systematically under-researched due to the increased difficulty of measuring uncertainty in regression models. Since normalizing flows offer a full predictive distribution instead of a point forecast, they facilitate direct usage of known heuristics for AL like Entropy or Least-Confident sampling. However, we show that most of these heuristics do not work well for normalizing flows in pool-based AL and we need more sophisticated algorithms to distinguish between aleatoric and epistemic uncertainty. In this work we propose BALSA, an adaptation of the BALD algorithm, tailored for regression with normalizing flows. With this work we extend current research on uncertainty quantification with normalizing flows \cite{berry2023normalizing, berry2023escaping} to real world data and pool-based AL with multiple acquisition functions and query sizes. We report SOTA results for BALSA across 4 different datasets and 2 different architectures.

主动学习归一化流不确定性估计回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。