改进音频分类主动学习策略,提升低预算下的标注效率。
Cover First, Disagree Softly: Rethinking Mismatch-First Active Learning for Frame-Level Audio Classification

- 用加权覆盖目标替代硬性分歧筛选,避免重复选择相似片段。
- 在两个数据集上均实现最优学习曲线面积,尤其在低预算下优势明显。
- 无需额外超参数,适配多类主动学习场景,适合资源受限研究者。
声事件检测依赖于昂贵的帧级强标签。主动学习通过选择最有助于分类器的音频段来缓解此问题。主流策略不匹配优先远端遍历(MFFT)结合两个分类器的分歧与所选片段的多样性,采用硬性顺序决策:先选高分歧段组,再以远端遍历分配剩余预算。我们在两个多标签数据集上发现该设计忽略所选片段间的相似性,在低预算下表现不佳,所有不匹配优先变体均低于其基础的几何策略。为此提出不匹配加权设施选址(MW-FL),以分歧加权覆盖目标使用全部预算,惩罚所选片段间的相似性。利用MFFT的分歧信号生成非负权重,无需引入超参数。在两种几何机制下,三种分歧使用方式的实验表明:所选片段覆盖度是主导因素,硬性分歧筛选有害,软性分歧加权可进一步提升性能。MW-FL在两个数据集上均取得最佳学习曲线面积。
原文摘要 · Abstract (English)
Sound event detection relies on frame-level strong labels whose annotation is expensive. Active learning addresses this problem by selecting the audio segments whose labels help the classifier most. One of the prevailing acquisition strategies for this task, mismatch-first farthest-traversal (MFFT), combines the disagreement between two classifiers and the diversity of the selected segments through hard sequential decisions. It selects whole groups of high-disagreement segments first and spreads only the remaining budget by farthest traversal. On two multi-label datasets we show that this design is blind to the similarity among the selected segments and fails under low budgets, with every mismatch-first variant ending below the plain geometric strategy it builds on. We propose mismatch-weighted facility location (MW-FL), which spends the entire budget through a disagreement-weighted coverage objective that penalizes similarity among the selected segments. The disagreement signal from MFFT is used to obtain the nonnegative weights of this facility-location objective, without introducing hyperparameters. Experiments across two geometric mechanisms with three ways of using disagreement show that coverage of the selected segments is the dominant factor, hard disagreement gating of selection is harmful on both mechanisms, and soft disagreement weighting helps on top of coverage. MW-FL attains the best area under the learning curve on both datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。