在设备端持续适应临床电话语音识别,解决真实场景下的识别误差问题。
Navigating the Reality Gap: On-Device Continual Adaptation of ASR for Clinical Telephony
- 采用参数高效方法与流式持续学习,在设备端实现语音识别模型动态优化。
- 模型在真实电话数据上错误率从11.59%升至41.71%,凸显现实差距。
- 负EWC系数可增强记忆回放效果,适合医疗等多领域实时语音场景。
自动语音识别(ASR)能显著减轻临床工作中的文书负担,但标准模型在真实电话场景中表现急剧下降,受限于噪声音频、方言差异及严格的数据本地化要求。本文使用Gram Vaani——一个涵盖农村医疗与农业热线的印地语电话语料库,作为临床语音在设备端约束下的最接近代理。结果显示,鲁棒的多语言模型IndicWav2Vec在标准清晰印地语上为11.59%的词错误率(WER),在该代理电话数据上增至41.71%。评估了从全量微调到参数高效LoRA及流式持续学习等多种设备端适应策略,覆盖多个基线、数据集与随机种子。聚焦持续学习,核心发现揭示经验回放(ER)与弹性权重固化(EWC,调节强度λ)间的关键交互:标准正EWC(λ>0)会抑制回放驱动的更新,限制适应;反转EWC强度(λ<0)表明其可在ER引导下充当方向性控制信号:负λ强化回放驱动的可塑性,而调度λ实现稳定性与可塑性的阶段调控。在多数据集验证中,多领域回放提供坚实基础,而EWC仅调节动态平衡而不改变最终性能。结果表明,有效设备端适应依赖于理解数据驱动与参数级学习信号的相互作用,而非孤立选择方法。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) can significantly reduce documentation burden in clinical workflows, but standard models degrade sharply in real-world telephony settings where noisy audio, dialectal variation, and strict data residency constraints prevent cloud-based adaptation. We study this "reality gap" using Gram Vaani: a telephonic Hindi corpus spanning rural healthcare and agricultural helplines, as the closest available proxy for clinical speech under strict on-device constraints. We show that a robust multilingual model (IndicWav2Vec) degrades from 11.59\% WER on standard clean Hindi to \textbf{41.71\% WER} on this proxy telephony data. We evaluate a progression of on-device adaptation regimes under realistic constraints, from full fine-tuning to parameter-efficient LoRA and stream-based continual learning, across multiple baselines, datasets, and seeds. Focusing on continual learning, our central finding highlights a critical interaction between Experience Replay (ER) and Elastic Weight Consolidation (EWC, parameterized by regularization strength $λ$). We show that standard positive EWC ($λ> 0$) can oppose replay-driven updates, limiting adaptation. Reversing EWC's strength ($λ< 0$) suggests that it can act as a directional control signal under ER-guided adaptation: negative $λ$ reinforces replay-driven plasticity, while a scheduled $λ$ enables phase-dependent control of stability and plasticity. Across evaluations on multiple datasets, we find that multi-domain replay provides a strong foundation for adaptation, while EWC modulates stability-plasticity dynamics without altering final performance. These results show that effective on-device adaptation depends on understanding how data-driven and parameter-level learning signals interact, rather than choosing methods in isolation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。