arXiv:2606.05571cs.SDeess.AS2026-06

统一音效数据集标签体系,实现跨库融合与可比分析

Sound Effects Dataset Unification With the Universal Category System

论文配图:Sound Effects Dataset Unification With the Universal Category System
图 1 · 摘自论文原文
  • 采用行业标准分级标签体系重构数据集标签
  • 构建含5.8万段音频的统一环境音数据集
  • 适合音效分类与生成研究者使用

音效数据集常使用不同的标签体系、分类结构和元数据格式,导致研究中分类与生成任务面临数据孤岛问题,难以比较结果或合并数据。本文提出一种模块化重标签框架,以行业标准的通用类别系统(UCS)作为统一结构基础。该开源框架通过基于规则的多阶段流水线与冲突解决机制,实现现有数据集标签向UCS的高自动转换率;提出分层数据集划分策略;并支持多数据集整合。为验证实用性,我们构建了公开的EnvSound-UCS数据集,包含来自AudioSet、FSD50K和ESC-50的58,057段环境音效。

原文摘要 · Abstract (English)

Sound effects (SFX) datasets and libraries often employ distinct tagging schemes, taxonomies, and metadata structures. This creates challenges for research on SFX classification and generation because incompatible taxonomies lead to siloed datasets that might require individualized approaches, result in non-comparable outcomes, and prevent data merging strategies. We propose a modular dataset relabeling framework that adopts the Universal Category System (UCS), an industry-standard hierarchical taxonomy for sound effects, as a shared structural foundation. This open-source framework enables us (i) to convert tags of existing datasets to UCS with a rule-based multi-stage pipeline and conflict resolution to achieve high automatic conversion rates, (ii) to suggest a stratified dataset split for the new labels, and (iii) to combine multiple datasets. To showcase the practical utility, we introduce the EnvSound-UCS dataset, a publicly available unified UCS-compliant dataset of environmental sounds with 58,057 sound clips from three sources: AudioSet, FSD50K, and ESC-50.

音效数据标签统一数据融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。