用Whisper模型一键识别野采集语音中的杂音、多说话人等噪声数据
Whilter: A Whisper-based Data Filter for "In-the-Wild" Speech Corpora Using Utterance-level Multi-Task Classification
- 基于Whisper编码器设计多任务分类器,同时判断五类语音异常
- 在三个子任务上F1超85%,误判率低至6.5%~7.8%
- 适合构建高质量语音数据集的研究者和工程师使用
近年来,大规模野外语音数据集因自监督学习需求增长而愈发普遍。这类数据常含多说话人、非目标语言及音乐等干扰因素,影响模型训练效果。本文提出Whilter模型,采用Whisper编码器与注意力分类器,一次性解决五类语音异常检测任务。研究还发布了两个主流野外语料库的标注子集。实验表明,Whilter在三个子任务上F1得分超过85%,等错误率降至6.5%~7.8%,优于当前最优的BEATs分类器,且处理速度显著快于多个单任务模型组合。
原文摘要 · Abstract (English)
Large-scale in-the-wild speech datasets have become more prevalent in recent years due to increased interest in models that can learn useful features from unlabelled data for tasks such as speech recognition or synthesis. These datasets often contain undesirable features, such as multiple speakers, non-target languages, and music, which may impact model learning. The Whilter model is proposed as a multitask solution to identify these undesirable samples. Whilter uses a Whisper encoder with an attention-based classifier to solve five diverse classification problems at once. In addition, an annotated dataset is published for a subset of two popular in-the-wild corpora. Whilter achieves F1 scores above 85% and equal error rates of 6.5% to 7.8% for three of five subtasks, outperforming a state-of-the-art BEATs classifier on speech-specific classes, with a notable decrease in processing time compared to a combination of single-task alternatives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。