arXiv:2509.26291cs.SDcs.AI2025-09

用自监督音频表征自动检测数据质量问题,省下大量人工审核

Representation-Based Data Quality Audits for Audio

  • 用音频表征排序识别数据错误、重复和标签问题
  • 在3个数据集上表现优于专门工具,减少标注工作量
  • 适合需要高质量音频数据的工业与研究团队

音频系统性能常受离题样本、近似重复和标签错误等问题制约。本文将图像领域的自清洁(SelfClean)表征排序审计框架迁移至音频领域,利用自监督音频表征统一识别常见数据质量缺陷,生成按严重程度排序的审查清单。该方法在ESC-50、GTZAN及一个工业级数据集上进行了测试,涵盖合成与真实噪声污染。结果表明,该框架在排名性能上达到当前最优,显著优于针对特定问题设计的基线方法,并能有效指导人工审核,大幅降低标注成本。

原文摘要 · Abstract (English)

Data quality issues such as off-topic samples, near duplicates, and label errors often limit the performance of audio-based systems. This paper addresses these issues by adapting SelfClean, a representation-to-rank data auditing framework, from the image to the audio domain. This approach leverages self-supervised audio representations to identify common data quality issues, creating ranked review lists that surface distinct issues within a single, unified process. The method is benchmarked on the ESC-50, GTZAN, and a proprietary industrial dataset, using both synthetic and naturally occurring corruptions. The results demonstrate that this framework achieves state-of-the-art ranking performance, often outperforming issue-specific baselines and enabling significant annotation savings by efficiently guiding human review.

音频数据质量审计自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。