arXiv:2501.03568cs.LGstat.ME2025-01综述被引 1

在标签成本高的情况下,高效进行两样本检验的实用方法。

Advanced Tutorial: Label-Efficient Two-Sample Tests

  • 结合主动学习思想,只标注关键特征以减少标签开销。
  • 保持统计有效性的同时,测试功效显著优于传统方法。
  • 适合临床研究、生物信息等标注代价高昂的场景。

假设检验是用于判断数据是否支持特定假设的统计推断方法。其中,两样本检验用于评估两组数据点是否来自相同分布,广泛应用于临床研究中比较治疗效果。本教程探讨在分析者拥有大量特征但确定这些特征的样本归属(标签)成本较高的情境下的两样本检验问题。该情形在机器学习中类似主动学习的研究背景。本教程将主动学习理念扩展至标签昂贵条件下的两样本检验,同时保证统计有效性与高检验功效。此外,还讨论了这些标签高效两样本检验的实际应用。

原文摘要 · Abstract (English)

Hypothesis testing is a statistical inference approach used to determine whether data supports a specific hypothesis. An important type is the two-sample test, which evaluates whether two sets of data points are from identical distributions. This test is widely used, such as by clinical researchers comparing treatment effectiveness. This tutorial explores two-sample testing in a context where an analyst has many features from two samples, but determining the sample membership (or labels) of these features is costly. In machine learning, a similar scenario is studied in active learning. This tutorial extends active learning concepts to two-sample testing within this \textit{label-costly} setting while maintaining statistical validity and high testing power. Additionally, the tutorial discusses practical applications of these label-efficient two-sample tests.

假设检验主动学习标签效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。