arXiv:2603.11168cs.LGcs.CL2026-03

用生物标志物监督提升亨廷顿病语音识别准确率

Huntington Disease Automatic Speech Recognition with Biomarker Supervision

  • 引入生物标志物辅助训练,引导模型关注疾病特异性语音特征
  • 识别错误模式随病情严重程度变化,非均匀改善,降低误识率至4.95%
  • 首个在高保真临床语料上系统评估多种语音识别模型的研究

针对亨廷顿病(HD)的病理语音自动语音识别(ASR)研究仍不充分,其不规则时序、不稳定发声和构音畸变给现有模型带来挑战。本文基于此前未用于端到端训练的高保真临床语音语料,系统比较多种ASR架构的性能,分析词错误率(WER)及替换、删除、插入错误模式。结果表明,不同模型在HD语音下呈现特定错误模式,其中Parakeet-TDT优于编码器-解码器与CTC基线模型。通过针对性适配,将WER从6.99%降至4.95%;同时提出基于生物标志物的辅助监督方法,发现错误行为随病情严重程度呈差异化重塑,而非统一改善。所有代码与模型已开源。

原文摘要 · Abstract (English)

Automatic speech recognition (ASR) for pathological speech remains underexplored, especially for Huntington's disease (HD), where irregular timing, unstable phonation, and articulatory distortion challenge current models. We present a systematic HD-ASR study using a high-fidelity clinical speech corpus not previously used for end-to-end ASR training. We compare multiple ASR families under a unified evaluation, analyzing WER as well as substitution, deletion, and insertion patterns. HD speech induces architecture-specific error regimes, with Parakeet-TDT outperforming encoder-decoder and CTC baselines. HD-specific adaptation reduces WER from 6.99% to 4.95% and we also propose a method for using biomarker-based auxiliary supervision and analyze how error behavior is reshaped in severity-dependent ways rather than uniformly improving WER. We open-source all code and models.

语音识别病理语音生物标志物亨廷顿病

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。