arXiv:2507.01372cs.CVcs.LG2025-07NeurIPS被引 4

用主动采样让AI高效估算大规模科学数据,人机协作减少人工成本。

Active Measurement: Efficient Estimation at Scale

  • AI预估数据,按重要性采样人类标注,动态优化模型
  • 即使模型不准也能精准估计,模型越准所需人工越少
  • 适用于高精度科学测量,适合追求效率与可靠性的研究者

人工智能有望通过极少的人工干预分析海量数据,推动科学发现。然而,当前工作流常缺乏足够的精度和统计保证。本文提出主动测量(Active Measurement),一种人机协同的科学测量框架:利用AI模型预测各单元的测量值,再通过重要性采样选取关键样本进行人工标注;每轮新标签更新模型,并同步改进无偏蒙特卡洛总和估计。该方法在模型不完美时仍能提供精确估计,在模型高度准确时仅需极少量人工。我们推导了新型估计器、加权方案与置信区间,实验证明其在多个测量任务中显著降低估计误差。

原文摘要 · Abstract (English)

AI has the potential to transform scientific discovery by analyzing vast datasets with little human effort. However, current workflows often do not provide the accuracy or statistical guarantees that are needed. We introduce active measurement, a human-in-the-loop AI framework for scientific measurement. An AI model is used to predict measurements for individual units, which are then sampled for human labeling using importance sampling. With each new set of human labels, the AI model is improved and an unbiased Monte Carlo estimate of the total measurement is refined. Active measurement can provide precise estimates even with an imperfect AI model, and requires little human effort when the AI model is very accurate. We derive novel estimators, weighting schemes, and confidence intervals, and show that active measurement reduces estimation error compared to alternatives in several measurement tasks.

主动学习科学计算高效估算人机协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。