用跨任务校准提升少标签场景下的推理精度
Prediction-Powered Inference Across Many Tasks for AI Evaluation & Social Science Research

- 基于相关任务的代理数据,通过跨任务校准优化推断
- 在标签稀缺时,置信区间宽度显著缩小,最高降37%
- 适合需要多任务验证的AI评估与社会科学研究
许多应用需在多个相关任务中进行统计有效推断,但每项假设仅能获取少量高质量标注。在人工智能评估中,这些任务可能对应模型在不同提示、子群体或假设下的行为;在社会科学调查中,则可能对应相关问题、人群或测量条件。预测驱动推断(PPI)利用大量廉价的代理测量来提升有限真实标签的推断效果,但传统方法独立处理各任务,未能利用相关任务间的共享结构,尤其在每任务标签极少时表现不佳。为此,我们提出一种多任务预测驱动推断框架,通过跨任务校准利用代理-真实关系中的共享结构,同时保留任务内校准与功效调优,构建准确的点估计与置信区间。理论证明:仅当代理-真实关系存在非线性结构时,效率提升才可能实现;仿射型跨任务校准在渐近意义下等价于使用原始代理。我们在合成与半合成数据集上及2024年美国大选语言模型审计案例研究中验证了该方法。基于大规模人工标注实验,结果显示,在标签稀缺时,跨任务校准可显著缩小置信区间宽度。
原文摘要 · Abstract (English)
Many applications require statistically valid inference across many related tasks, while using only a handful of high-quality labels per hypothesis. In AI evaluation, these tasks may correspond to model behaviors across prompts, subgroups, or hypotheses; in social science surveys, they may correspond to related questions, populations, or measurement conditions. Prediction-powered inference (PPI) uses abundant but inexpensive proxy measurements to improve inference from limited, ground-truth labels, but commonly used methods treat tasks independently and therefore fail to exploit shared structure across related tasks. This limitation is especially important in settings where only a small number of labels are available per task. To address this issue, we introduce a multi-task prediction-powered inference framework that uses labeled data from related tasks to improve power while preserving task-specific inference. Our methods exploit the shared structure in the proxy-ground-truth relationship through cross-task recalibration, while retaining within-task rectification and power tuning to construct accurate point estimates and confidence intervals. We prove that efficiency gains beyond power-tuned PPI are only possible when the proxy-ground-truth relationship contains nonlinear structure; affine cross-task recalibrations are asymptotically equivalent to using the original proxy. We complement our theoretical findings with experiments on synthetic and semi-synthetic datasets, as well as a case study auditing language models on election-related information during the 2024 U.S. presidential election. Using a large human-annotation study, we show that cross-task recalibration can substantially reduce confidence interval widths when labels are scarce.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。