arXiv:2602.19578stat.MLcs.LG2026-02

不依赖贝叶斯后验,用反曲率实现不确定性感知的数据选择

Goal-Oriented Influence-Maximizing Data Acquisition for Learning and Optimization

  • 基于目标函数梯度与损失曲率的联合优化,选择对目标影响最大的样本
  • 在图像、文本分类及超参调优任务中,用更少标注达到相同性能
  • 适合需高效标注的深度学习训练与黑箱优化场景

主动数据获取在深度神经网络的学习与优化中至关重要,但现有方法多依赖难以可靠获取的预测不确定性。为此,我们提出面向目标的影响最大化数据获取(GOIMDA),该算法无需显式后验推断,通过逆曲率保持不确定性感知。GOIMDA 通过最大化候选输入对用户指定目标函数(如测试损失、预测熵或优化器推荐设计值)的期望影响来选择样本。利用一阶影响函数,我们推导出一个可计算的获取规则,融合目标梯度、训练损失曲率与样本对模型参数的敏感性。理论证明,在广义线性模型下,GOIMDA 在修正目标对齐与预测偏差后,近似于预测熵最小化,从而在不维护贝叶斯后验的前提下实现不确定性感知。实验表明,在图像与文本分类、噪声全局优化基准及神经网络超参数调优等任务中,GOIMDA 均以显著更少的标注样本或函数评估次数达成目标性能,优于基于不确定性的主动学习与高斯过程贝叶斯优化基线。

原文摘要 · Abstract (English)

Active data acquisition is central to many learning and optimization tasks in deep neural networks, yet remains challenging because most approaches rely on predictive uncertainty estimates that are difficult to obtain reliably. To this end, we propose Goal-Oriented Influence- Maximizing Data Acquisition (GOIMDA), an active acquisition algorithm that avoids explicit posterior inference while remaining uncertainty-aware through inverse curvature. GOIMDA selects inputs by maximizing their expected influence on a user-specified goal functional, such as test loss, predictive entropy, or the value of an optimizer-recommended design. Leveraging first-order influence functions, we derive a tractable acquisition rule that combines the goal gradient, training-loss curvature, and candidate sensitivity to model parameters. We show theoretically that, for generalized linear models, GOIMDA approximates predictive-entropy minimization up to a correction term accounting for goal alignment and prediction bias, thereby, yielding uncertainty-aware behavior without maintaining a Bayesian posterior. Empirically, across learning tasks (including image and text classification) and optimization tasks (including noisy global optimization benchmarks and neural-network hyperparameter tuning), GOIMDA consistently reaches target performance with substantially fewer labeled samples or function evaluations than uncertainty-based active learning and Gaussian-process Bayesian optimization baselines.

主动学习数据获取优化不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。