arXiv:2603.27508cs.SD2026-03

测试语音模型在运动后说话时的鲁棒性,发现不同模型表现差异大。

Investigation on the Robustness of Acoustic Foundation Models on Post Exercise Speech

  • 统一评测框架下比较多种语音模型在运动后语音的表现。
  • 运动后语音错误率最高达14.57% WER,非流利说话者更难识别。
  • 模型微调能显著提升性能,但效果不稳定,需关注说话流畅度影响。

自动语音识别(ASR)在平静状态下的语音上研究广泛,但在运动后生理变化下的鲁棒性仍待探索。与静息语音相比,运动后语音常出现微呼吸、非语义停顿、发声不稳和重复等现象,导致转录难度增加。本文在统一评估协议下,对多个声学基础模型在运动后语音上的表现进行基准测试。比较了序列到序列模型(Whisper 和 FunASR/Paraformer)以及自监督编码器结合CTC解码的模型(Wav2Vec2、HuBERT、WavLM),涵盖开箱即用推理和运动后领域内微调两种场景。在 Static/Post-All 基准上,多数模型在运动后语音上性能下降;其中 FunASR 在 Post-All 上表现出最强基线鲁棒性,达到 14.57% WER 与 8.21% CER。微调显著提升了多个 CTC 模型的性能,而 Whisper 展现出不稳定的适应能力。作为探索性案例研究,我们按流利程度分层分析结果:尽管非流利子集较小,但始终比流利子集更具挑战性。总体而言,研究发现运动后语音识别的鲁棒性高度依赖模型,领域内适配可大幅提升性能但并非普遍稳定,未来研究应明确区分语言流利度与运动引起的语音变化。

原文摘要 · Abstract (English)

Automatic speech recognition (ASR) has been extensively studied on neutral and stationary speech, yet its robustness under post-exercise physiological shift remains underexplored. Compared with resting speech, post-exercise speech often contains micro-breaths, non-semantic pauses, unstable phonation, and repetitions caused by reduced breath support, making transcription more difficult. In this work, we benchmark acoustic foundation models on post-exercise speech under a unified evaluation protocol. We compare sequence-to-sequence models (Whisper and FunASR/Paraformer) and self-supervised encoders with CTC decoding (Wav2Vec2, HuBERT, and WavLM), under both off-the-shelf inference and post-exercise in-domain fine-tuning. Across the Static/Post-All benchmark, most models degrade on post-exercise speech, while FunASR shows the strongest baseline robustness at 14.57% WER and 8.21% CER on Post-All. Fine-tuning substantially improves several CTC-based models, whereas Whisper shows unstable adaptation. As an exploratory case study, we further stratify results by fluent and non-fluent speakers; although the non-fluent subset is small, it is consistently more challenging than the fluent subset. Overall, our findings show that post-exercise ASR robustness is strongly model-dependent, that in-domain adaptation can be highly effective but not uniformly stable, and that future post-exercise ASR studies should explicitly separate fluency-related effects from exercise-induced speech variation.

语音识别运动后语音模型鲁棒性语音质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。