统一建模说话人分离与语音识别,提升复杂对话的转录准确率
SoulX-Transcriber: A Robust End-to-End Framework for Multi-Speaker Speech Transcription

- 基于大语言模型联合建模说话人分离与语音识别
- 在多个公开数据集上实现高精度多说话人转录
- 适合需要处理真实对话场景的语音应用开发者
近年来,自动语音识别(ASR)和大语言模型(LLM)的发展显著提升了语音理解能力。然而,多说话人语音转录仍面临巨大挑战,受限于说话人声音高度相似、快速发言切换、语音重叠以及说话人边界分割不准确等问题。这些问题在真实对话音频中尤为突出,因说话人动态和声学条件变化剧烈。本文提出SoulX-Transcriber,一个统一的多说话人转录系统,将说话人分离(SD)与ASR联合建模于基于LLM的框架中。该方法采用两阶段训练策略:第一阶段通过面向说话人的多任务连续预训练增强说话人表征学习与边界感知;第二阶段通过监督微调进一步优化模型,在复杂多说话人条件下实现准确的端到端说话人标注转录。SoulX-Transcriber在AliMeeting、AISHELL-4和AMI等多个公开基准上表现优异,且具备良好的跨领域适应性。
原文摘要 · Abstract (English)
Recent advances in Automatic Speech Recognition (ASR) and Large Language Models (LLMs) have significantly improved speech understanding capabilities. However, multi-speaker speech transcription remains challenging task, constrained by highly similar speaker voices, rapid turn-taking transitions, overlapping utterances and inaccurate speaker boundary segmentation. These challenges become particularly pronounced in real-world conversational audio, where speaker dynamics and acoustic conditions are highly variable. This technical report presents SoulX-Transcriber, a unified multi-speaker transcription system that jointly models speaker diarization (SD) and ASR within an LLM-based framework. SoulX-Transcriber adopts a two-stage training strategy to improve both speaker discrimination and transcription robustness. In the first stage, speaker-aware multi-task continuous pre-training enhances speaker representation learning and boundary perception. In the second stage, supervised fine-tuning further optimizes the model for accurate end-to-end speaker-attributed transcription under complex multi-speaker conditions. SoulX-Transcriber delivers strong performance and robustness across multiple public benchmarks, including AliMeeting, AISHELL-4, and AMI, while maintaining high adaptability to multi-domain scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。