用冻结的多语言语音识别模型实现高效跨语种说话人归属
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
- 用弱标签训练说话人模块,不修改原有语音识别模型
- 仅用单语数据训练,在多语种重叠语音上表现良好
- 无需微调主模型,适合快速部署于实际场景
说话人归属自动语音识别(SA-ASR)旨在转录语音并准确分配发言者身份。现有方法通常依赖复杂的模块化系统或需对联合模块进行大量微调,限制了其适应性和整体效率。本文提出一种新方法,利用冻结的多语言语音识别模型,在不修改模型的前提下,通过标准单语语音识别数据集实现说话人归属。该方法仅需训练一个说话人模块,基于弱标签预测说话人嵌入。尽管训练数据为非重叠单语数据,模型仍能在多种多语言数据集(包括存在重叠语音的场景)中有效提取说话人属性。实验表明,性能媲美强基线,展现出良好的鲁棒性与实际应用潜力。
原文摘要 · Abstract (English)
Speaker-attributed automatic speech recognition (SA-ASR) aims to transcribe speech while assigning transcripts to the corresponding speakers accurately. Existing methods often rely on complex modular systems or require extensive fine-tuning of joint modules, limiting their adaptability and general efficiency. This paper introduces a novel approach, leveraging a frozen multilingual ASR model to incorporate speaker attribution into the transcriptions, using only standard monolingual ASR datasets. Our method involves training a speaker module to predict speaker embeddings based on weak labels without requiring additional ASR model modifications. Despite being trained exclusively with non-overlapping monolingual data, our approach effectively extracts speaker attributes across diverse multilingual datasets, including those with overlapping speech. Experimental results demonstrate competitive performance compared to strong baselines, highlighting the model's robustness and potential for practical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。