构建泰国远场会议语音数据集,提升中文以外语言的语音识别鲁棒性。
LOTUSDIS: A Thai far-field meeting corpus for robust conversational ASR
- 采集114小时泰国多人自发对话,9种麦克风远距离(0.12-10米)独立录音
- 远场语音识别错误率从81.6降至49.5,微调后整体错误率降为38.3
- 公开数据与基线系统,助力泰国语远场语音识别研究
我们提出LOTUSDIS,一个公开可用的泰语会议语音数据集,旨在推动远场对话式语音识别的发展。该数据集包含114小时自发、非脚本的对话,每场持续15-20分钟,参与者为三人,重叠说话频繁且自然。语音由九个独立的单通道设备同步录制,覆盖六种麦克风类型,距离范围为0.12米至10米,真实保留混响、噪声和设备失真效应,无需使用麦克风阵列。我们提供标准训练、开发和测试划分,并发布可复现的基线系统。在零样本和微调条件下评估多个Whisper变体。未经微调的模型随距离增加性能显著下降,表明预训练数据与泰语远场语音存在不匹配。在LOTUSDIS上微调后,泰国Whisper基线的整体词错误率(WER)从64.3降至38.3,远场WER从81.6降至49.5,尤其在最远距离麦克风上提升显著。结果凸显了多样化距离训练数据对鲁棒语音识别的重要性。数据集遵循CC-BY-SA 4.0许可,同时发布训练与评估脚本,以促进该领域的可复现研究。
原文摘要 · Abstract (English)
We present LOTUSDIS, a publicly available Thai meeting corpus designed to advance far-field conversational ASR. The dataset comprises 114 hours of spontaneous, unscripted dialogue collected in 15-20 minute sessions with three participants, where overlapping speech is frequent and natural. Speech was recorded simultaneously by nine independent single-channel devices spanning six microphone types at distances from 0.12 m to 10 m, preserving the authentic effects of reverberation, noise, and device coloration without relying on microphone arrays. We provide standard train, dev, test splits and release a reproducible baseline system. We benchmarked several Whisper variants under zero-shot and fine-tuned conditions. Off-the-shelf models showed strong degradation with distance, confirming a mismatch between pre-training data and Thai far-field speech. Fine-tuning on LOTUSDIS dramatically improved robustness: a Thai Whisper baseline reduced overall WER from 64.3 to 38.3 and far-field WER from 81.6 to 49.5, with especially large gains on the most distant microphones. These results underscore the importance of distance-diverse training data for robust ASR. The corpus is available under CC-BY-SA 4.0. We also release training and evaluation scripts as a baseline system to promote reproducible research in this field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。