提出鲁棒主动采样方法,确保不比随机采样差,且在不确定性估计准确时表现更优。
Robust Sampling for Active Statistical Inference
- 根据不确定性评分质量动态融合均匀采样与主动采样
- 在真实社会科学研究数据集上验证,始终优于或等同于均匀采样
- 适合对采样可靠性要求高、但不确定模型置信度的场景
主动统计推断是一种结合AI辅助数据收集的新方法。给定可标注数据点数量的预算,并假设可访问一个预测模型,其核心思想是优先标注模型最不确定的数据以提升估计精度。然而,当不确定性估计不准确时,主动采样可能导致噪声极大,甚至劣于简单的均匀采样。本文提出鲁棒采样策略,确保最终估计器性能从不劣于均匀采样。当不确定性估计可靠时,通常显著优于标准主动推断。该方法通过最优地在均匀采样与主动采样间插值,结合鲁棒优化思想实现。我们在一系列来自计算社会科学和调查研究的真实数据集上验证了该方法的有效性。
原文摘要 · Abstract (English)
Active statistical inference is a new method for inference with AI-assisted data collection. Given a budget on the number of labeled data points that can be collected and assuming access to an AI predictive model, the basic idea is to improve estimation accuracy by prioritizing the collection of labels where the model is most uncertain. The drawback, however, is that inaccurate uncertainty estimates can make active sampling produce highly noisy results, potentially worse than those from naive uniform sampling. In this work, we present robust sampling strategies for active statistical inference. Robust sampling ensures that the resulting estimator is never worse than the estimator using uniform sampling. Furthermore, with reliable uncertainty estimates, the estimator usually outperforms standard active inference. This is achieved by optimally interpolating between uniform and active sampling, depending on the quality of the uncertainty scores, and by using ideas from robust optimization. We demonstrate the utility of the method on a series of real datasets from computational social science and survey research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。