arXiv:2505.24224eess.AS2025-05中稿 · Interspeech 2025被引 4

用提示专家混合模型,让语音识别实时适配老人说话特点。

MOPSA: Mixture of Prompt-Experts Based Speaker Adaptation for Elderly Speech Recognition

  • 用聚类提示构建专家库,动态路由选择适配
  • 在两个老年语音数据集上降低0.86%的词错误率
  • 适合需要快速适配老人语音的智能助手场景

本文提出一种基于提示专家混合的说话人自适应方法(MOPSA),用于老年人语音识别。该方法可在无训练样本的情况下实现零样本、实时适配未见说话人,并融合针对老年人的领域知识。通过K均值聚类提取前K个最具区分度的说话人提示簇作为专家,由路由器网络动态组合这些提示专家。针对Whisper模型,分别使用声学和语言级编码器与解码器提示来建模老年人之间的变异性。在英语DementiaBank Pitt和粤语JCCOCC MoCA老年语音数据集上的实验表明,线上MOPSA适配相比说话人无关(SI)模型,在词错误率(WER)和字符错误率(CER)上分别绝对降低0.86%和1.47%(相对降低4.21%和5.40%),具有统计显著性。相比离线批处理适配,其实时因子(RTF)提速最高达16.12倍。

原文摘要 · Abstract (English)

This paper proposes a novel Mixture of Prompt-Experts based Speaker Adaptation approach (MOPSA) for elderly speech recognition. It allows zero-shot, real-time adaptation to unseen speakers, and leverages domain knowledge tailored to elderly speakers. Top-K most distinctive speaker prompt clusters derived using K-means serve as experts. A router network is trained to dynamically combine clustered prompt-experts. Acoustic and language level variability among elderly speakers are modelled using separate encoder and decoder prompts for Whisper. Experiments on the English DementiaBank Pitt and Cantonese JCCOCC MoCA elderly speech datasets suggest that online MOPSA adaptation outperforms the speaker-independent (SI) model by statistically significant word error rate (WER) or character error rate (CER) reductions of 0.86% and 1.47% absolute (4.21% and 5.40% relative). Real-time factor (RTF) speed-up ratios of up to 16.12 times are obtained over offline batch-mode adaptation.

语音识别老年语音提示学习实时适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。