arXiv:2608.01865cs.CL2026-08

分析失语症语音对模型内部表示的影响,找出适配优化的关键层。

Analyzing Speech Condition Effects in Dysarthric ASR: A Layer-wise Probing Study

  • 分层探测发现失语语音在深层仍难提取音位信息。
  • 中层微调仅用0.65%参数,使识别率提升2.89%。
  • 适合低资源失语症语音识别的高效适配研究者。

自动语音识别(ASR)在失语症语音上性能急剧下降,但其对模型内部表示的影响尚不明确。本文针对普通话失语症语音,在三种匹配转录条件(原始失语语音、说话人条件零样本语音合成、无条件语音合成)下,对Transformer ASR编码器进行分层探测分析。结果显示:音位边界信息在所有条件下各层均弱;音位身份在合成语音的深层可恢复,但在失语语音中依然较差;识别困难集中于最深层。此外,词汇声调是所有条件下持续的错误来源。基于此,采用层选择性LoRA微调发现,中层适配(第7层或第5-8层)可在仅训练0.16%和0.65%适配器参数的情况下,分别实现6.67%和2.89%的相对性能提升。而上层适配对合成语音更有效。研究将表示分析与参数高效微调结合,为低资源普通话失语症语音识别提供层感知优化方案。

原文摘要 · Abstract (English)

Automatic speech recognition (ASR) performance degrades sharply on dysarthric speech, yet how disordered articulation reshapes a model's internal representations is underexplored. We conduct a layer-wise probing analysis of a transformer ASR encoder on Mandarin dysarthric speech under three transcript-matched conditions: original dysarthric speech, speaker conditioned zero-shot TTS resynthesis, and unconditioned TTS. Probing reveals a task- and condition-dependent representation hierarchy: phoneme boundary information remains weak across all layers for dysarthric speech; phoneme identity is recoverable in deep layers for synthetic speech, but remains poor for dysarthric speech; and recognition difficulty is concentrated in the deepest layers. Furthermore, lexical tone is a persistent error source across all conditions. Guided by these insights, layer-selective LoRA shows that mid-layer adaptation (layer 7 or layers 5-8) recovers near-full encoder performance on dysarthric speech within 6.67% and 2.89% relative margins while training only 0.16% and 0.65% of adapter parameters. Conversely, upper-layer adaptation benefits synthetic speech more than dysarthric speech. These findings link representation analysis to parameter-efficient fine-tuning and motivate layer-aware adaptation for low-resource Mandarin dysarthric ASR.

语音识别失语症微调分层分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。