arXiv:2511.13529cs.CLcs.AI2025-11被引 3

发布匈牙利语对话语音数据集,推动低资源语言语音识别发展

Toward Conversational Hungarian Speech Recognition: Introducing the BEA-Large and BEA-Dialogue Datasets

  • 构建了255小时自发语音和85小时对话数据,含精细元信息
  • 在自发语音上实现14.18%词错误率,重复语音达4.8%
  • 适合研究对话语音识别与说话人分离,为其他语言提供范式

自动语音识别(ASR)的进步主要依赖高资源语言的海量数据,而匈牙利语等语言因缺乏自然对话语料仍处于欠代表状态。为此,我们从此前未处理的匈牙利语语音语料BEA中构建了两个新数据集——BEA-Large和BEA-Dialogue。BEA-Large在BEA-Base基础上扩展了255小时来自433名说话人的自发语音,并附带详细的段级元数据;BEA-Dialogue包含85小时自然对话,划分为独立说话人子集,支持对话式ASR与说话人分离研究。我们使用公开ASR模型建立了可复现基线,微调后的Fast Conformer模型在自发语音上达到14.18%词错误率,在重复语音上达4.8%。说话人分离实验的错误率在12.46%至17.40%之间,为后续改进提供参考。结果表明,对话式ASR仍面临不流畅、重叠和非正式表达等挑战。通过发布这些数据集与基线,我们旨在推动匈牙利语语音技术发展,并为其他语言构建自发与对话基准提供方法框架。

原文摘要 · Abstract (English)

The advancement of automatic speech recognition (ASR) has been largely enhanced by extensive datasets in high-resource languages, while languages such as Hungarian remain underrepresented due to limited spontaneous and conversational corpora. To address this gap, we introduce two new datasets -- BEA-Large and BEA-Dialogue -- constructed from the previously unprocessed portions of the Hungarian speech corpus named BEA. BEA-Large extends BEA-Base with 255 hours of spontaneous speech from 433 speakers, enriched with detailed segment-level metadata. BEA-Dialogue, comprising 85 hours of spontaneous conversations, is a Hungarian speech corpus featuring natural dialogues partitioned into speaker-independent subsets, supporting research in conversational ASR and speaker diarization. We establish reproducible baselines on these datasets using publicly available ASR models, with the fine-tuned Fast Conformer model achieving word error rates as low as 14.18% on spontaneous and 4.8% on repeated speech. Diarization experiments yield diarization error rates between 12.46% and 17.40%, providing reference points for future improvements. The results highlight the persistent difficulty of conversational ASR, particularly due to disfluencies, overlaps, and informal speech patterns. By releasing these datasets and baselines, we aim to advance Hungarian speech technology and offer a methodological framework for developing spontaneous and conversational benchmarks in other languages.

语音识别匈牙利语对话数据集说话人分离

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。