arXiv:2606.30237cs.CL2026-06

对比人与自动系统识别荷兰口吃语音,发现个性化模型已超越人类听者。

Comparing Human and Automatic Recognition of Dutch Dysarthric Continuous Speech: A Case Study

  • 用单个严重口吃者数据,对比人类与三个主流语音识别系统表现
  • 所有系统与人类平均词错误率超70%,微调后降至23%以上仍较高
  • 个性化模型优于人类听者,对日常交流有实用潜力

为开发个性化口吃语音识别(DSR)模型,本研究对比了人类听者与三种先进离线语音识别系统(Whisper-large-V3、Google Chirp 3、Omnilingual)在单个严重口吃者读诵和自发连续语音上的识别表现。结果表明,人类听者与三类系统平均词错误率(WER)均超过70%,显示口吃语音识别极具挑战性。在口吃语音上进行微调显著降低了WER。尽管总体WER仍高于23%,个性化DSR模型性能已超越人类听者,且正逐渐接近支持日常交流的实用水平。未来研究应聚焦于自发语音及更长语句的个性化DSR提升,尤其关注特定音素识别。

原文摘要 · Abstract (English)

In our goal to develop personalised dysarthric speech recognition (DSR) models, this study compared the recognition performances of human listeners and those of three state-of-the-art, off-the-shelf ASR systems (Whisper-large-V3, Google Chirp 3, and Omnilingual) on the recognition of Dutch continuous read and spontaneous speech from a single speaker with severe dysarthria. Results showed that both humans listeners and the three off-the-shelf ASR systems exhibit word error rates (WER) exceeding 70% on average, indicating that DSR is highly challenging for both humans and ASR systems. Fine-tuning on the dysarthric speech significantly reduced WER. Although overall WERs are still quite high (>23%), the personalised DSR models outperformed the human listeners, and performance is getting closer to being useful for supporting day-to-day communication of dysarthric speakers. Future research should focus on improving personalized DSR on spontaneous speech and longer utterances in the case of read speech, with a specific focus on particular phonemes.

语音识别口吃语音个性化模型自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。