在有限标注预算下,智能分配真实标签与偏好判断资源,提升模型学习效率。
Labels or Preferences? Budget-Constrained Learning with Human Judgments over AI-Generated Outputs
- 构建单调缺失数据框架,优化真实标签与偏好判断的预算分配
- 提出PCAL方法,显著降低估计方差,理论证明渐近最优性
- 适用于多种任务,特别适合标注成本高的现代AI场景
随着人类偏好反馈在评估AI生成伪标签中的应用日益广泛,如何在有限预算下制定合理的数据采集策略成为关键挑战。本文针对固定标注预算下真实标签与成对偏好之间的最优分配问题,提出基于半参数推断的解决方案。通过将预算分配建模为单调缺失数据框架,我们设计了偏好校准主动学习(PCAL)方法,能够学习最优数据采集策略并构建统计高效的分布函数估计器。理论上,我们证明了所提估计器的渐近最优性,并建立了关键鲁棒性保证,确保即使在非参数模型估计不准确时仍保持稳定性能。该框架具有高度灵活性,直接优化估计器方差,无需闭式解即可适用广泛问题。实验结果表明,模拟与真实数据均验证了本方法的优越性与实用性。
原文摘要 · Abstract (English)
The increasing reliance on human preference feedback to judge AI-generated pseudo labels has created a pressing need for principled, budget-conscious data acquisition strategies. We address the crucial question of how to optimally allocate a fixed annotation budget between ground-truth labels and pairwise preferences in AI. Our solution, grounded in semi-parametric inference, casts the budget allocation problem as a monotone missing data framework. Building on this formulation, we introduce Preference-Calibrated Active Learning (PCAL), a novel method that learns the optimal data acquisition strategy and develops a statistically efficient estimator for functionals of the data distribution. Theoretically, we prove the asymptotic optimality of our PCAL estimator and establish a key robustness guarantee that ensures robust performance even with poorly estimated nuisance models. Our flexible framework applies to a general class of problems, by directly optimizing the estimator's variance instead of requiring a closed-form solution. This work provides a principled and statistically efficient approach for budget-constrained learning in modern AI. Simulations and real-data analysis demonstrate the practical benefits and superior performance of our proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。