arXiv:2504.18582cs.SDcs.CL2025-04被引 5

通过微调Wav2Vec提升库尔德语语音说话人分离效果

Speaker Diarization for Low-Resource Languages Through Wav2vec Fine-Tuning

  • 在库尔德语语料上微调Wav2Vec 2.0模型,迁移多语言表征
  • 相比基线方法,错误率降低7.2%,聚类纯度提升13%
  • 适合低资源语言语音处理、多语种会议系统等场景

说话人分离是语音处理中的基础任务,旨在按说话人划分音频流。尽管现有先进模型在高资源语言上表现优异,但库尔德语等低资源语言因标注数据少、方言多样及频繁语码转换而面临独特挑战。本文通过在专用库尔德语语料上训练Wav2Vec 2.0自监督学习模型,利用迁移学习将其他语言中习得的多语言表征适配至库尔德语语音的音系与声学特征。相较于基线方法,本方法使说话人分离错误率降低7.2%,聚类纯度提升13%。研究结果表明,对现有模型进行优化可显著提升低资源语言的分离性能。该工作为开发库尔德语媒体转录服务及多语种呼叫中心、视频会议系统的说话人分割技术提供了实用价值,并为其他未充分研究语言的语音系统建设奠定基础,推动语音技术的公平性。

原文摘要 · Abstract (English)

Speaker diarization is a fundamental task in speech processing that involves dividing an audio stream by speaker. Although state-of-the-art models have advanced performance in high-resource languages, low-resource languages such as Kurdish pose unique challenges due to limited annotated data, multiple dialects and frequent code-switching. In this study, we address these issues by training the Wav2Vec 2.0 self-supervised learning model on a dedicated Kurdish corpus. By leveraging transfer learning, we adapted multilingual representations learned from other languages to capture the phonetic and acoustic characteristics of Kurdish speech. Relative to a baseline method, our approach reduced the diarization error rate by seven point two percent and improved cluster purity by thirteen percent. These findings demonstrate that enhancements to existing models can significantly improve diarization performance for under-resourced languages. Our work has practical implications for developing transcription services for Kurdish-language media and for speaker segmentation in multilingual call centers, teleconferencing and video-conferencing systems. The results establish a foundation for building effective diarization systems in other understudied languages, contributing to greater equity in speech technology.

说话人分离低资源语言Wav2Vec语音处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。