arXiv:2508.06372cs.SDcs.AI2025-08AAAI被引 29

用统一模型同时解决说话人分离与识别,支持灵活注册。

SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language Models

  • 端到端联合建模说话人分离与语音识别
  • 在多个数据集上超越现有分步系统表现
  • 支持不同注册人数和场景的鲁棒识别

说话人分离与识别(SDR)旨在预测音频片段中“谁在何时说了什么”,是会议转录和对话系统等多说话人场景的关键任务。现有系统多采用分步框架,结合说话人分离(SD)与自动语音识别(ASR)等多个模块,存在错误传播、重叠语音处理难、以及SD与ASR任务缺乏联合优化等问题。为此,我们提出SpeakerLM,一种用于SDR的统一多模态大语言模型,以端到端方式联合执行SD与ASR。为适配多样化实际场景,我们在SpeakerLM中引入灵活的说话人注册机制,支持不同注册设置下的SDR。SpeakerLM通过大规模真实数据的多阶段训练逐步构建。大量实验表明,SpeakerLM展现出强大的数据扩展能力与泛化性能,在域内与域外公共基准测试中均优于当前最优分步基线。此外,实验结果验证了所提注册机制能有效保障SpeakerLM在不同注册条件及注册人数变化下的稳定表现。

原文摘要 · Abstract (English)

The Speaker Diarization and Recognition (SDR) task aims to predict "who spoke when and what" within an audio clip, which is a crucial task in various real-world multi-speaker scenarios such as meeting transcription and dialogue systems. Existing SDR systems typically adopt a cascaded framework, combining multiple modules such as speaker diarization (SD) and automatic speech recognition (ASR). The cascaded systems suffer from several limitations, such as error propagation, difficulty in handling overlapping speech, and lack of joint optimization for exploring the synergy between SD and ASR tasks. To address these limitations, we introduce SpeakerLM, a unified multimodal large language model for SDR that jointly performs SD and ASR in an end-to-end manner. Moreover, to facilitate diverse real-world scenarios, we incorporate a flexible speaker registration mechanism into SpeakerLM, enabling SDR under different speaker registration settings. SpeakerLM is progressively developed with a multi-stage training strategy on large-scale real data. Extensive experiments show that SpeakerLM demonstrates strong data scaling capability and generalizability, outperforming state-of-the-art cascaded baselines on both in-domain and out-of-domain public SDR benchmarks. Furthermore, experimental results show that the proposed speaker registration mechanism effectively ensures robust SDR performance of SpeakerLM across diverse speaker registration conditions and varying numbers of registered speakers.

说话人识别端到端大模型语音分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。