arXiv:2504.12279cs.SDcs.CL2025-04被引 2

用几何变换修复失语语音,让语音识别更鲁棒。

Dysarthria Normalization via Local Lie Group Transformations for Robust ASR

  • 将语音畸变建模为频谱图的局部李群变换,通过神经网络反推变形场
  • 在真实失语语音上实现最高17个百分点的错误率降低,且对正常语音无损
  • 无需真实病理语音训练,适合医疗语音识别场景

我们提出一种基于几何的方法,通过将时间、频率和振幅畸变建模为频谱图的平滑局部李群变换来归一化失语语音。标量场通过指数映射生成这些形变,一个神经网络仅使用合成扭曲的健康语音进行训练,以在测试时推断场并应用近似逆变换。我们引入自发对称性破缺(SSB)势能,促使模型发现非平凡的场配置。在真实病理语音上,系统表现稳定:在具有挑战性的TORGO语句上最多降低17个百分点的词错误率(WER),WER方差下降16%,且在干净的CommonVoice数据上无性能退化。字符和音素错误率同步改善,验证了语言相关性。结果表明,几何结构化的形变可为失语语音识别提供一致的零样本鲁棒性提升。

原文摘要 · Abstract (English)

We present a geometry-driven method for normalizing dysarthric speech by modeling time, frequency, and amplitude distortions as smooth, local Lie group transformations of spectrograms. Scalar fields generate these deformations via exponential maps, and a neural network is trained - using only synthetically warped healthy speech - to infer the fields and apply an approximate inverse at test time. We introduce a spontaneous-symmetry-breaking (SSB) potential that encourages the model to discover non-trivial field configurations. On real pathological speech, the system delivers consistent gains: up to 17 percentage-point WER reduction on challenging TORGO utterances and a 16 percent drop in WER variance, with no degradation on clean CommonVoice data. Character and phoneme error rates improve in parallel, confirming linguistic relevance. Our results demonstrate that geometrically structured warping provides consistent, zero-shot robustness gains for dysarthric ASR.

语音识别失语矫正几何建模鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。