用文本嵌入选例,让大模型无微调就能听懂难说话。
TICL: Text-Embedding KNN For Speech In-Context Learning Unlocks Speech Recognition Abilities of Large Multimodal Models
- 基于文本语义相似度挑选上下文示例
- 在多种难语音上相对错误率降低84.7%
- 无需微调,适合快速部署到多场景
语音基础模型近期展现出语音上下文学习(SICL)能力。有效选择上下文示例对SICL性能至关重要,但相关方法仍不充分。本文提出文本嵌入KNN用于SICL(TICL),一种简单流程,利用语义上下文提升现成大视觉-语言模型的语音识别能力,无需微调。在包括口音英语、多语言语音和儿童语音在内的多个挑战性任务中,该方法使模型性能超越零样本基准,相对词错误率(WER)最高降低84.7%。我们通过消融实验验证了方法的鲁棒性与高效性。
原文摘要 · Abstract (English)
Speech foundation models have recently demonstrated the ability to perform Speech In-Context Learning (SICL). Selecting effective in-context examples is crucial for SICL performance, yet selection methodologies remain underexplored. In this work, we propose Text-Embedding KNN for SICL (TICL), a simple pipeline that uses semantic context to enhance off-the-shelf large multimodal models' speech recognition ability without fine-tuning. Across challenging automatic speech recognition tasks, including accented English, multilingual speech, and children's speech, our method enables models to surpass zero-shot performance with up to 84.7% relative WER reduction. We conduct ablation studies to show the robustness and efficiency of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。