新基准显示弱监督在真实任务中表现远超预期,需上千标注样本才可媲美。
Stronger Than You Think: Benchmarking Weak Supervision on Realistic Tasks
- 构建真实场景弱监督基准BOXWRENCH,涵盖高类别复杂性与领域知识
- 在多数任务中,弱监督仅需少量标注数据即超越监督学习
- 适合研究弱监督机制或数据效率的开发者参考
弱监督(WS)是一种通过多种噪声但低成本的弱标签来源自动标注训练数据的高效学习方法。尽管广泛应用,其实际价值难以评估,因配置涉及数据源、标注函数(LFs)、聚合模型(称作标签模型)及下游模型流程等多个可调参数。现有评测工具通常局限在特定组件或特殊场景,且常使用简化任务或编写不佳的默认标注函数,导致结果无法推广至真实环境。为此,我们提出新基准BOXWRENCH,更贴近真实应用:包含高类别数量与分布不均的任务、需领域知识的任务,并支持跨多语言语料复用标注函数。所有标注函数均按真实场景精心设计。相比现有基准,我们发现,在多数情况下,监督学习需1000个以上标注样本才能达到弱监督性能。
原文摘要 · Abstract (English)
Weak supervision (WS) is a popular approach for label-efficient learning, leveraging diverse sources of noisy but inexpensive weak labels to automatically annotate training data. Despite its wide usage, WS and its practical value are challenging to benchmark due to the many knobs in its setup, including: data sources, labeling functions (LFs), aggregation techniques (called label models), and end model pipelines. Existing evaluation suites tend to be limited, focusing on particular components or specialized use cases. Moreover, they often involve simplistic benchmark tasks or de-facto LF sets that are suboptimally written, producing insights that may not generalize to real-world settings. We address these limitations by introducing a new benchmark, BOXWRENCH, designed to more accurately reflect real-world usages of WS. This benchmark features tasks with (1) higher class cardinality and imbalance, (2) notable domain expertise requirements, and (3) opportunities to re-use LFs across parallel multilingual corpora. For all tasks, LFs are written using a careful procedure aimed at mimicking real-world settings. In contrast to existing WS benchmarks, we show that supervised learning requires substantial amounts (1000+) of labeled examples to match WS in many settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。