自动生成高质量弱标签,兼顾覆盖范围与可靠性。
Structured Exploration and Exploitation of Label Functions for Automated Data Annotation

- 从表面、结构、语义多层探索标签规则,提升多样性。
- 实现最高98.9%标签覆盖率,弱标签质量提升87%。
- 适合需要高效标注的机器学习研究与应用者。
高质量标注数据对训练可靠机器学习模型至关重要,但人工标注成本高且易出错。程序化标注通过标签函数(LFs)——即自动生成弱标签的启发式规则——缓解该问题。然而,现有自动化LF生成方法要么依赖大语言模型生成表层启发式规则,要么基于手工设计的基元进行建模合成,常导致覆盖范围有限且标签质量不可靠。本文提出EXPONA,一种系统化平衡多样性和可靠性的自动化程序化标注框架。EXPONA从表面、结构和语义多层次系统探索标签函数,并采用可靠性感知机制抑制噪声或冗余规则,同时保留互补信号。我们在11个跨领域分类数据集上进行广泛实验,结果表明EXPONA持续优于当前最优自动化LF生成方法:标签覆盖率高达98.9%,弱标签质量提升达87%,下游加权F1提升最高46%。这表明,多层级探索与可靠性过滤的结合,有效提升了不同任务中标签质量和下游性能的一致性。
原文摘要 · Abstract (English)
High-quality labeled data is critical for training reliable machine learning and deep learning models, yet manual annotation remains costly and error-prone. Programmatic labeling addresses this challenge by using label functions (LFs), i.e., heuristic rules that automatically generate weak labels for training datasets. However, existing automated LF generation methods either rely on large language models (LLMs) to synthesize surface-level heuristics or employ model-based synthesis over hand-crafted primitives. These approaches often result in limited coverage and unreliable label quality. In this paper, we introduce EXPONA, an automated framework for programmatic labeling that formulates LF generation as a principled process balancing diversity and reliability. EXPONA systematically explores multi-level LFs, spanning surface, structural, and semantic perspectives. EXPONA further applies reliability-aware mechanisms to suppress noisy or redundant heuristics while preserving complementary signals. To evaluate EXPONA, we conducted extensive experiments on eleven classification datasets across diverse domains. Experimental results show that EXPONA consistently outperformed state-of-the-art automated LF generation methods. Specifically, EXPONA achieved nearly complete label coverage (up to 98.9%), improved weak label quality by up to 87%, and yielded downstream performance gains of up to 46% in weighted F1. These results indicate that EXPONA's combination of multi-level LF exploration and reliability-aware filtering enabled more consistent label quality and downstream performance across diverse tasks by balancing coverage and precision in the generated LF set.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。