构建多口音长对话数据集,评估语音识别在真实客服场景的泛化能力。
AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR
- 基于14种英语口音的自发角色扮演对话,覆盖16类服务场景
- 不同口音和分段方式下识别准确率差异显著,通用美式基准不具普适性
- 专为评测设计,无公开数据重叠,适合研究跨口音鲁棒性
评估面向对话式AI的英语语音识别系统仍具挑战性,因现有公开语料库多为短片段、朗读或准备好的语音,且缺乏明确方言标注,难以评估对多样化用户的鲁棒性。本文提出AppTek Call-Center Dialogues语料库,包含14种英语口音的自发角色扮演客服对话,覆盖16类服务场景。该数据集专为评测设计,发布前音频与文本均未公开,避免与大规模预训练语料重叠。我们测试了多种开源语音识别系统在不同分段策略下的表现。结果表明,不同口音与分段方法间存在显著性能差异,说明在通用美式英语基准上表现良好并不意味着能推广至其他口音。
原文摘要 · Abstract (English)
Evaluating English ASR systems for conversational AI applications remains difficult, as many publicly available corpora are either pre-segmented into short segments, consist of read or prepared speech, or lack explicit dialect annotations to evaluate robustness for a diverse user base. This work presents the AppTek Call-Center Dialogues corpus, a collection of spontaneous, role-played agent-customer conversations spanning fourteen English accents covering sixteen service-oriented scenarios. The dataset was commissioned specifically for evaluation and none of the audio or text was publicly available prior to release, reducing the risk of overlap with existing large-scale pretraining corpora. We benchmark a set of open-source ASR systems under different segmentation approaches. Results show substantial variation across accents and segmentation methods, indicating that good performance on general American English benchmarks does not necessarily generalize to other accents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。