arXiv:2605.17860cs.CLcs.AI2026-05中稿 · and presented at S…

构建首个多口音NLP学术讨论语音数据集,助力提升真实场景下语音识别鲁棒性。

PAREDA: A Multi-Accent Speech Dataset of Natural Language Processing Research Discussions

论文配图:PAREDA: A Multi-Accent Speech Dataset of Natural Language Processing Research Discussions
图 1 · 摘自论文原文
  • 采集澳、印、中三种英语口音的NLP论文讨论对话,含即兴陈述与问答场景。
  • 零样本测试下模型词错误率高,表明数据具有挑战性;微调后WER显著下降。
  • 适合研究语音识别鲁棒性、跨口音适配及专业领域语音系统开发的团队。

尽管现代自动语音识别(ASR)系统在基准语料上表现良好,但在面对真实世界中的口音、即兴表达和领域特定语言时性能常下降。本文提出首个多口音语音数据集PAREDA,包含来自澳大利亚、印度英语和中文英语口音的研究者对自然语言处理论文的学术讨论。每段录音包含一段即兴独白(论文摘要)和一段问答对话,涵盖大量技术术语与口语现象。我们评估了当前最优ASR模型在该数据集上的表现,分析口音混合与语速加快的影响。结果显示,在零样本设置下,模型表现较差,证实数据集的挑战性;但通过在PAREDA上微调,词错误率(WER)明显降低,说明该数据集捕捉到现有语料中缺失的语言特征。PAREDA为构建和评估更具鲁棒性与包容性的专用领域语音识别系统提供了宝贵资源。

原文摘要 · Abstract (English)

While modern Automatic Speech Recognition (ASR) systems achieve high accuracy on benchmark corpora, their performance often degrades when there is real-world variability. This work focuses on variability arising due to accented, spontaneous, and domain-specific speech. In particular, we introduce PAper REading DAtaset (PAREDA), a first-of-its-kind multi-accent speech dataset consisting of discussions on academic Natural Language Processing (NLP) papers between speakers with Australian, Indian-English, and Chinese English accents. Each session elicits a spontaneous monologue (a summary of a paper's abstract) and a non-monologue (a question-and-answer session between participants), resulting in a corpus rich with technical jargon and conversational phenomena. We evaluate the performance of SOTA ASR models on PAREDA, analysing the impact of accent mixing and increased speech rate. Our results show that, in the zero-shot setting, models perform worse, confirming the dataset's challenging nature. However, fine-tuning on PAREDA significantly reduces the Word Error Rate (WER), demonstrating that our dataset captures linguistic characteristics often missing from existing corpora. PAREDA serves as a valuable new resource for building and evaluating more robust and inclusive ASR systems for specialised, real-world applications.

语音识别多口音NLP数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。