自动标注音频数据,提升特定场景下分类准确率。
TriA Pipeline: A Large-Scale Automatic Audio Annotation Pipeline For Audio Classification In Specific Scenarios

- 用自动管道将多场景音频转为带事件标注的高质量训练数据
- 构建2130小时、431类的音频数据集,提升分类性能3.97%以上
- 适合需要大量标注数据的智能家庭、语音识别等场景
针对多数应用场景(如家庭环境)中音频标注数据稀缺的问题,本文提出一种自动音频标注流水线——TriA Pipeline,可高效将各类场景下的音频转化为带有音频事件标注的高质量训练数据。基于该流水线构建了包含超过2130小时音频、覆盖431个音频类别的TriA数据集,并从中提取出基于先验知识引导的子集(TriA$_{\mathrm{GK}}$),在三个家庭场景音频分类任务上进行对比实验。结果表明,仅使用人工标注数据时,加入TriA$_{\mathrm{GK}}$后平均准确率提升3.97%,宏平均F1值提升3.35%,验证了TriA$_{\mathrm{GK}}$及整个流水线的有效性。
原文摘要 · Abstract (English)
There are some datasets of varying scales for audio classification (AC) applied to different tasks. However, annotated data is limited for most scenarios, such as domestic environments. To address this challenge, we propose an $\textbf{A}$utomatic $\textbf{A}$udio $\textbf{A}$nnotation Pipeline--TriA Pipeline, which can efficiently convert audio from various scenarios into high-quality training data with audio event annotations. A TriA dataset was constructed with the TriA Pipeline, over 2130 hours of audio covering 431 audio classes. Furthermore, we partitioned a prior-knowledge-guided subset (TriA$_{\mathrm{GK}}$) from TriA and conduct comparative experiments on three domestic AC tasks. Comparing the result on manually annotated data only and that on manually annotated data combines TriA$_{\mathrm{GK}}$, TriA$_{\mathrm{GK}}$ could achieve average relative gains of 3.97% in accuracy and 3.35% in Macro-F1, validating the effectiveness of TriA$_{\mathrm{GK}}$ and the TriA Pipeline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。