用加权方法纠正音频分类评估中的抽样偏差,提升真实性能预测准确性。
Sampling Bias Compensation for Robust Evaluation of Audio Classification Systems with Partially Labeled Evaluation Datasets
- 通过特征空间的密度比估计,对小样本标注集进行重要性加权补偿。
- 在音频场景分类任务中,加权后准确率更接近全量数据真实表现。
- 适用于标注预算有限、需高精度评估的音频系统部署场景。
声学机器学习系统的性能通常在完全标注的测试集上评估。然而在实际部署中,对持续采集的大规模音频数据进行彻底标注往往不可行,因此性能评估通常依赖于可用数据中的小规模标注子集,这引入了抽样偏差,可能严重扭曲评估指标。本文研究在严格标注预算约束下,补偿评估子集偏差的方法。我们探讨重要性加权技术是否能通过补偿选择偏差来缓解这一问题。具体地,实现并对比了三种密度比估计方法:核密度估计(KDE)、逻辑回归和k近邻(kNN),利用部署音频的特征空间表示。为模拟真实部署场景,采用基于主动学习的五种不同采样策略生成标注子集。在音频场景分类(ASC)基准上的实验表明,重要性加权能持续提供更真实的准确率估计,显著缩小子集指标与真实评估性能之间的差距。
原文摘要 · Abstract (English)
The performance of acoustic machine learning systems is commonly evaluated using fully annotated test sets. In real-world deployments, however, exhaustively labeling large volumes of continuously collected audio data is often infeasible. Consequently, performance assessment typically relies on a small labeled subset of the available data, introducing a sampling bias that can severely distort evaluation metrics. This paper studies methods for compensating the bias in evaluation-labeled subsets under strict annotation-budget constraints. We study whether importance weighting techniques can mitigate this discrepancy by compensating for the selection bias. Specifically, we implement and compare three density-ratio estimation methods: kernel density estimation (KDE), logistic regression, and k-nearest neighbors (kNN), utilizing feature-space representations of the deployed audio. To emulate realistic deployment scenarios, the labeled subsets are generated using five distinct sampling strategies based on active learning techniques. Experiments conducted on an audio scene classification (ASC) benchmark demonstrate that importance weighting consistently yields more realistic accuracy estimates, significantly reducing the gap between subset-based metrics and the true evaluation performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。