arXiv:2508.03896stat.MLcs.LG2025-08TPAMI被引 1

为弱监督标签概率提供置信区间,提升预测可靠性。

Reliable Programmatic Weak Supervision with Confidence Intervals for Label Probabilities

  • 构建不确定集封装多种弱标注函数的任意行为与类型
  • 在多个基准数据集上显著优于现有方法,标签概率更可靠
  • 适合需要可信标签概率的场景,如医疗、金融领域

准确标注数据集既耗时又昂贵。针对未标注数据集,程序化弱监督通过多个弱标注函数(LFs)获取标签的概率预测,这些函数仅提供粗略标签猜测。然而,弱标注函数常具有不同类型的输出和未知的相互依赖关系,导致预测不可靠。此外,现有程序化弱监督技术无法评估标签概率预测的可靠性。本文提出一种新方法,可为标签概率提供置信区间,并获得更可靠的预测。所提方法使用不确定性分布集合,包容弱标注函数的任意行为与类型。在多个基准数据集上的实验表明,该方法优于当前最优技术,且所生成的置信区间具有实际应用价值。

原文摘要 · Abstract (English)

The accurate labeling of datasets is often both costly and time-consuming. Given an unlabeled dataset, programmatic weak supervision obtains probabilistic predictions for the labels by leveraging multiple weak labeling functions (LFs) that provide rough guesses for labels. Weak LFs commonly provide guesses with assorted types and unknown interdependences that can result in unreliable predictions. Furthermore, existing techniques for programmatic weak supervision cannot provide assessments for the reliability of the probabilistic predictions for labels. This paper presents a methodology for programmatic weak supervision that can provide confidence intervals for label probabilities and obtain more reliable predictions. In particular, the methods proposed use uncertainty sets of distributions that encapsulate the information provided by LFs with unrestricted behavior and typology. Experiments on multiple benchmark datasets show the improvement of the presented methods over the state-of-the-art and the practicality of the confidence intervals presented.

弱监督置信区间标签质量不确定性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。