arXiv:2502.09363cs.LG2025-02

研究固定时长标签对事件识别准确率的影响,揭示其代价并提出更优标注策略。

The Accuracy Cost of Weakness: A Theoretical Analysis of Fixed-Segment Weak Labeling for Events in Time

  • 分析固定时长片段标注的准确性与标注成本关系
  • 证明真实事件激活的最优标注方法在多数场景更优
  • 为自适应弱标注提供理论支持,适合序列标注任务优化

准确标签对构建稳健机器学习模型至关重要。本文建模了一种常见弱标注过程:标注者对固定长度数据段是否包含某类事件进行存在性标注,例如音频中是否存在鸟鸣事件。若某段充分覆盖该类事件,则标记为“存在”。我们分析了片段长度如何影响标签准确率及所需标注数量,并将此固定长度标注方法与使用真实事件激活构建片段的“理想”标注方法(记为oracle)进行比较。结果表明,在大多数现实场景中,oracle方法在准确率和标注成本上均优于固定长度方法。研究为模仿oracle过程的自适应弱标注策略提供了理论依据,也为序列标注任务中的弱标注优化奠定了基础。

原文摘要 · Abstract (English)

Accurate labels are critical for deriving robust machine learning models. Labels are used to train supervised learning models and to evaluate most machine learning paradigms. In this paper, we model the accuracy and cost of a common weak labeling process where annotators assign presence or absence labels to fixed-length data segments for a given event class. The annotator labels a segment as "present" if it sufficiently covers an event from that class, e.g., a birdsong sound event in audio data. We analyze how the segment length affects the label accuracy and the required number of annotations, and compare this fixed-length labeling approach with an oracle method that uses the true event activations to construct the segments. Furthermore, we quantify the gap between these methods and verify that in most realistic scenarios the oracle method is better than the fixed-length labeling method in both accuracy and cost. Our findings provide a theoretical justification for adaptive weak labeling strategies that mimic the oracle process, and a foundation for optimizing weak labeling processes in sequence labeling tasks.

弱监督序列标注标注效率理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。