arXiv:2601.13458stat.MLcs.AI2026-01被引 2

在有限标注预算下,智能分配真实标签与偏好判断资源,提升模型学习效率。

Labels or Preferences? Budget-Constrained Learning with Human Judgments over AI-Generated Outputs

  • 构建单调缺失数据框架,优化真实标签与偏好判断的预算分配
  • 提出PCAL方法,显著降低估计方差,理论证明渐近最优性
  • 适用于多种任务,特别适合标注成本高的现代AI场景

随着人类偏好反馈在评估AI生成伪标签中的应用日益广泛,如何在有限预算下制定合理的数据采集策略成为关键挑战。本文针对固定标注预算下真实标签与成对偏好之间的最优分配问题,提出基于半参数推断的解决方案。通过将预算分配建模为单调缺失数据框架,我们设计了偏好校准主动学习(PCAL)方法,能够学习最优数据采集策略并构建统计高效的分布函数估计器。理论上,我们证明了所提估计器的渐近最优性,并建立了关键鲁棒性保证,确保即使在非参数模型估计不准确时仍保持稳定性能。该框架具有高度灵活性,直接优化估计器方差,无需闭式解即可适用广泛问题。实验结果表明,模拟与真实数据均验证了本方法的优越性与实用性。

原文摘要 · Abstract (English)

The increasing reliance on human preference feedback to judge AI-generated pseudo labels has created a pressing need for principled, budget-conscious data acquisition strategies. We address the crucial question of how to optimally allocate a fixed annotation budget between ground-truth labels and pairwise preferences in AI. Our solution, grounded in semi-parametric inference, casts the budget allocation problem as a monotone missing data framework. Building on this formulation, we introduce Preference-Calibrated Active Learning (PCAL), a novel method that learns the optimal data acquisition strategy and develops a statistically efficient estimator for functionals of the data distribution. Theoretically, we prove the asymptotic optimality of our PCAL estimator and establish a key robustness guarantee that ensures robust performance even with poorly estimated nuisance models. Our flexible framework applies to a general class of problems, by directly optimizing the estimator's variance instead of requiring a closed-form solution. This work provides a principled and statistically efficient approach for budget-constrained learning in modern AI. Simulations and real-data analysis demonstrate the practical benefits and superior performance of our proposed method.

主动学习偏好学习预算优化半参数推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。