arXiv:2601.13319cs.CL2026-01ACL被引 1

构建标准化框架,统一14种方言阿拉伯语语音识别数据集

Arab Voices: Mapping Standard and Dialectal Arabic Speech Technology

  • 通过计算分析方言特征与音质指标,系统评估31个数据集
  • 发现方言信号强弱不一、录音条件差异大,现有数据难以直接比较
  • 推出Arab Voices框架,支持可复现的现代方言语音识别评测

方言阿拉伯语(DA)语音数据在领域覆盖、方言标注方式和录制条件上差异显著,导致跨数据集比较与模型评估困难。为刻画这一现状,我们对广泛使用的DA语料库训练集进行计算分析,评估语言层面的‘方言度’及音频质量客观指标。结果表明,各数据集在声学条件和方言信号强度与一致性方面存在显著异质性,凸显了超越粗粒度标签的标准化表征需求。为减少碎片化并支持可复现评估,我们提出Arab Voices——一个标准化的阿拉伯语方言语音识别框架。该框架统一提供涵盖14种方言的31个数据集,配备统一元数据与评估工具。我们进一步对多种近期语音识别系统进行基准测试,建立现代方言语音识别的强基线。

原文摘要 · Abstract (English)

Dialectal Arabic (DA) speech data vary widely in domain coverage, dialect labeling practices, and recording conditions, complicating cross-dataset comparison and model evaluation. To characterize this landscape, we conduct a computational analysis of linguistic ``dialectness'' alongside objective proxies of audio quality on the training splits of widely used DA corpora. We find substantial heterogeneity both in acoustic conditions and in the strength and consistency of dialectal signals across datasets, underscoring the need for standardized characterization beyond coarse labels. To reduce fragmentation and support reproducible evaluation, we introduce Arab Voices, a standardized framework for DA ASR. Arab Voices provides unified access to 31 datasets spanning 14 dialects, with harmonized metadata and evaluation utilities. We further benchmark a range of recent ASR systems, establishing strong baselines for modern DA ASR.

语音识别阿拉伯语方言处理数据标准化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。