arXiv:2606.16539eess.AScs.SD2026-06中稿 · Interspeech 2026

用语音文本提示实现老人语音实时适配,零样本下显著降低识别错误。

Decoding while Adapting: Zero-Shot Online Speaker Adaptation via Audio-Textual Prompts for Elderly Speech Recognition

论文配图:Decoding while Adapting: Zero-Shot Online Speaker Adaptation via Audio-Textual Prompts for Elderly Speech Recognition
图 1 · 摘自论文原文
  • 通过融合当前及前几句语音文本嵌入生成紧凑说话人提示。
  • 在两个老年语音数据集上,WER和CER分别降低0.61%和1.22%。
  • 适合需要实时适配陌生老人语音的智能助老系统使用。

本文提出一种基于跨话语语音-文本提示的老年人语音识别说话人适配方法,实现零样本、实时适配未见说话人。从当前及若干前序话语中提取语音与文本嵌入,以跨模态方式融合生成更一致的紧凑说话人提示,优于i-vector和ECAPA-TDNN特征。在英文DementiaBank Pitt与粤语JCCOCC MoCA老年语音数据集上的实验表明,该在线适配方法相比说话人无关(SI)模型,在词错误率(WER)或字符错误率(CER)上实现了0.61%和1.22%的绝对降低(相对降低2.99%和4.48%),且实时因子(RTF)速度提升最高达9.83倍,优于离线批处理适配。

原文摘要 · Abstract (English)

This paper proposes a novel cross-utterance audio-textual prompts based speaker adaptation approach for elderly speech recognition. It enables zero-shot, real-time adaptation to unseen speakers. Speech and text embeddings are extracted from the current and a few preceding utterances, before being fused in a cross-modal manner to produce compact speaker prompts that are more consistent than i/x-vectors and ECAPA-TDNN features. Experiments on the English DementiaBank Pitt and Cantonese JCCOCC MoCA elderly speech datasets suggest that the proposed online adaptation outperforms the speaker-independent (SI) model by statistically significant word error rate (WER) or character error rate (CER) reductions of 0.61% and 1.22% absolute (2.99% and 4.48% relative). Real-time factor (RTF) speed-up ratios of up to 9.83 times are obtained over offline batch-mode adaptation.

语音识别老年语音实时适配跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。