构建真实鸟类声音的细粒度多标签数据集,用于评估弱监督学习性能。
Merlin L48 Spectrogram Dataset
- 基于真实鸟类叫声构建细粒度多标签数据集,模拟真实场景的单正例标注。
- 在真实数据上测试现有方法,发现性能显著低于合成数据上的表现。
- 提供带领域先验的扩展设置,适合研究弱监督与领域知识融合的方法。
在单正例多标签(SPML)设置中,每张图像仅标注一个类别存在,其余类别的真实状态未知。挑战在于缩小此类部分标注设置与全监督学习之间的性能差距,而后者通常需要大量标注成本。以往的SPML方法多在合成数据集上评估,这些数据通过从如Pascal VOC、COCO、NUS-WIDE和CUB200等全标注数据集中随机抽取单一正标签生成。然而,这种合成方式无法反映真实世界场景,也难以捕捉导致困难误分类的细微复杂性。本文提出L48数据集,一个源自真实鸟类声音录音的细粒度、真实世界的多标签数据集。该数据集提供了自然的SPML设置,包含单正例标注,在具有挑战性的细粒度领域中进行评估,并进一步提供两种扩展设置,其中引入领域先验以获得额外负标签。我们在L48上对现有SPML方法进行了基准测试,观察到其在真实数据上的表现与合成数据上存在显著差异,并分析了方法的弱点,凸显了更真实、更具挑战性的基准的必要性。
原文摘要 · Abstract (English)
In the single-positive multi-label (SPML) setting, each image in a dataset is labeled with the presence of a single class, while the true presence of other classes remains unknown. The challenge is to narrow the performance gap between this partially-labeled setting and fully-supervised learning, which often requires a significant annotation budget. Prior SPML methods were developed and benchmarked on synthetic datasets created by randomly sampling single positive labels from fully-annotated datasets like Pascal VOC, COCO, NUS-WIDE, and CUB200. However, this synthetic approach does not reflect real-world scenarios and fails to capture the fine-grained complexities that can lead to difficult misclassifications. In this work, we introduce the L48 dataset, a fine-grained, real-world multi-label dataset derived from recordings of bird sounds. L48 provides a natural SPML setting with single-positive annotations on a challenging, fine-grained domain, as well as two extended settings in which domain priors give access to additional negative labels. We benchmark existing SPML methods on L48 and observe significant performance differences compared to synthetic datasets and analyze method weaknesses, underscoring the need for more realistic and difficult benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。