用合成语音提升英文说话人分离模型对印尼语对话的适应性
Domain Adaptation of the Pyannote Diarization Pipeline for Conversational Indonesian Audio
- 用神经文本转语音生成合成数据,适配低资源语言
- 25小时合成数据使错误率降至29.24%,比基线降13.68%
- 适合做小语种语音分析或语音分离系统迁移的研究者
本研究提出针对印尼语对话音频的说话人分离领域自适应方法。针对英语主导的分离模型在低资源语言上表现差的问题,采用神经文本转语音技术生成合成数据进行训练。实验使用小规模数据集(171样本)和大规模数据集(25小时合成语音)。结果表明,基于AMI语料库训练的基准模型在零样本迁移至印尼语时,说话人分离错误率(DER)为53.47%。经领域自适应后,小数据模型将DER降至34.31%(1轮训练)和34.81%(2轮训练);25小时数据训练模型达到最优性能,DER为29.24%,较基线降低13.68%,同时保持99.06%召回率和87.14%F1分数。
原文摘要 · Abstract (English)
This study presents a domain adaptation approach for speaker diarization targeting conversational Indonesian audio. We address the challenge of adapting an English-centric diarization pipeline to a low-resource language by employing synthetic data generation using neural Text-to-Speech technology. Experiments were conducted with varying training configurations, a small dataset (171 samples) and a large dataset containing 25 hours of synthetic speech. Results demonstrate that the baseline \texttt{pyannote/segmentation-3.0} model, trained on the AMI Corpus, achieves a Diarization Error Rate (DER) of 53.47\% when applied zero-shot to Indonesian. Domain adaptation significantly improves performance, with the small dataset models reducing DER to 34.31\% (1 epoch) and 34.81\% (2 epochs). The model trained on the 25-hour dataset achieves the best performance with a DER of 29.24\%, representing a 13.68\% absolute improvement over the baseline while maintaining 99.06\% Recall and 87.14\% F1-Score.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。