arXiv:2409.11746cs.SDeess.AS2024-09被引 5

构建统一音频事件标签体系,解决多数据集标签不兼容问题

SALT: Standardized Audio event Label Taxonomy

  • 基于AudioSet层级结构,标准化24个环境音数据集标签
  • 实现跨数据集标签映射,支持统一分类与数据整合
  • 提供便捷工具包,助力跨源音频实验与探索

机器听觉系统常依赖固定分类体系来组织和标注音频数据,这对训练和评估深度神经网络(DNN)等监督算法至关重要。然而,现有分类体系受限于特定应用的预定义类别,难以融入新声音,且因标签标准不一,导致跨数据集兼容性差。为此,本文提出SALT:标准化音频事件标签体系。该体系在AudioSet本体的层次结构基础上,扩展并标准化了24个公开环境音数据集的标签,使不同数据集的类别标签可映射至统一系统。我们还开发了一个新的Python工具包,支持标签的导航与使用,便于跨数据集标签搜索与层次化探索。该工具包可轻松实现多源数据聚合,显著降低跨数据集实验的门槛。

原文摘要 · Abstract (English)

Machine listening systems often rely on fixed taxonomies to organize and label audio data, key for training and evaluating deep neural networks (DNNs) and other supervised algorithms. However, such taxonomies face significant constraints: they are composed of application-dependent predefined categories, which hinders the integration of new or varied sounds, and exhibits limited cross-dataset compatibility due to inconsistent labeling standards. To overcome these limitations, we introduce SALT: Standardized Audio event Label Taxonomy. Building upon the hierarchical structure of AudioSet's ontology, our taxonomy extends and standardizes labels across 24 publicly available environmental sound datasets, allowing the mapping of class labels from diverse datasets to a unified system. Our proposal comes with a new Python package designed for navigating and utilizing this taxonomy, easing cross-dataset label searching and hierarchical exploration. Notably, our package allows effortless data aggregation from diverse sources, hence easy experimentation with combined datasets.

音频识别标签标准化数据融合工具包

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。