用未标注数据的预测增强假设检验,提升统计功效。
Semi-Supervised Hypothesis Testing by Betting on Predictions
- 基于预测构建赌注型e统计量,实现任意时间有效的序列检验
- 在标签漂移或概念漂移下,即使预测不准也能保持非平凡检验力
- 适用于低相关性场景,特别适合大模型评估等弱监督任务
我们提出一种基于赌注的假设检验框架,利用未标注数据的预测来增强序列假设检验的效能。在仅有少量联合分布$(X,Y)$样本,但有额外边缘分布$X$的未标注样本条件下,研究如何利用未标注数据推断$Y$的分布及$Yig|X$的条件分布。引入e统计量并构建序列检验,在标准分布假设(标签漂移或概念漂移)下,证明该检验具有任意时间有效性。进一步表明,对于二元数据,该e统计量具有非平凡的检验力。关键优势在于,即使基础预测不准确,仍能保持上述性质。通过模拟和大语言模型评估应用,验证了相较于基线方法(包括预测驱动推断)的显著性能提升,且在未标注数据有限、预测相关性较弱时仍有效。
原文摘要 · Abstract (English)
We introduce a testing-by-betting framework that leverages predictions on unlabeled data to enhance the power of sequential hypothesis testing. Given limited samples from the joint distribution of $(X,Y)$, and additional unlabeled samples from the marginal of $X$, we ask how unlabeled data can be used to hypothesize about the distribution of $Y$, and the conditional distribution of $Y\mid X$. We introduce an e-statistic and use it to construct a sequential test. Under standard distributional assumptions -- label shift or concept shift -- we establish that the test is anytime valid. Furthermore, we show that for binary data, the e-statistic has non-trivial power. Crucially, our approach retains these properties even when the underlying predictions are inaccurate. Through simulations and applications to large language models evaluation, we demonstrate power gains over baseline approaches, including prediction-powered inference. These gains persist even with relatively limited unlabeled data and when predictions have low accuracy due to weak correlation between $X$ and $Y$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。