用FiLM技术让语音模型适应病理性发音,不改主干也能提升识别准确率。
FiLM-Based Speaker Conditioning of a SpeechLLM for Pathological Speech Recognition

- 通过FiLM在每层注入语音特征向量,动态调节模型表示以适配特定患者。
- 在西班牙语和英语病理性语音上达到与微调相当的识别性能。
- 保持原模型对正常语音的识别能力,适合临床多病种场景使用。
自动语音识别(ASR)在标准语音上已取得显著进展,但来自神经疾病患者的病理性语音仍是重大挑战。本文研究基于特征逐元素线性调制(FiLM)的说话人条件化方法,将x向量提取的信息注入冻结的ASR编码器每个Transformer层中,以自适应调整内部表示,无需修改基础模型权重。我们在西班牙语和英语病理性语音数据集上,对比了该方法与标准微调及参数高效微调基线的ASR性能,并结合后处理进行评估。此外,还检验了适配后模型在回答语音相关问题上的能力。结果表明,说话人条件化后的ASR性能可媲美已有适配策略,同时保留对非条件化语音的识别能力。
原文摘要 · Abstract (English)
Automatic speech recognition (ASR) has advanced remarkably for standard speech; however, pathological speech from neurological conditions remains a significant challenge. We investigate speaker conditioning via Feature-wise Linear Modulation (FiLM), injecting x-vector-derived information into each transformer layer of a frozen ASR encoder to adapt internal representations to individual pathological speakers without modifying base model weights. We benchmark this for the ASR task against standard and parameter-efficient fine-tuning baselines, complemented by post-processing, on Spanish and English pathological speech. Additionally, we evaluate if the adapted model preserves the ability to answer speech-related questions. Results show that speaker-conditioned ASR is competitive with established adaptation strategies while retaining performance on non-conditioned speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。