基于语音大模型的端到端多说话人识别与转录系统,可精准分辨谁在何时说了什么。
DiaScriber: A Speech LLM for Joint Diarization and Transcription in Multi-Speaker Scenarios
- 构建多样化数据流,模拟复杂多说话人场景及重叠语音。
- 在多个测试集上超越现有方法,未见场景下仍表现优异。
- 适合需要高精度多说话人语音分析的场景,如会议记录、访谈转写。
多说话人自动语音识别(MSASR)旨在联合预测内容转录、说话人身份和时间戳,解决‘谁在何时说了什么’这一关键问题,在真实多说话人场景中具有重要应用价值。然而,面对快速话轮切换、重叠语音以及复杂多变的多说话人场景,当前方法仍面临显著挑战。本文提出DiaScriber,一个基于语音大语言模型的端到端多说话人分离与转录模型。我们首先构建多样化的数据流水线,覆盖各类复杂多说话人场景,包括验证与修正、话轮切换与重叠语音模拟、多模态标注等。DiaScriber基于预训练的Qwen3.5-Omni模型,采用三阶段训练策略:持续预训练、监督微调和强化学习。实验表明,DiaScriber在多个大规模多说话人场景测试集上性能优于对比方法,并在未见过的复杂场景中展现出出色的泛化能力。
原文摘要 · Abstract (English)
Multi-speaker automatic speech recognition (MSASR) aims to jointly predict content transcriptions, speaker identities, and timestamps, thereby addressing the key question of "who spoke what and when" and holds substantial practical value in real-world multi-speaker scenarios. However, MSASR still encounters considerable challenges in the presence of fast turn transitions, overlapping speech, and complex, diverse multi-speaker scenarios. In this work, we propose DiaScriber, an end-to-end multi-speaker diarization and transcription model built on a speech large language model. We first construct diverse data pipelines to cover a wide variety of multi-speaker scenarios and their complexities, including validation and refinement, turn-transition and overlapping-speech simulation, and multimodal annotation. Furthermore, DiaScriber is developed based on the pretrained version of Qwen3.5-Omni through a three-stage training strategy involving continual pretraining, supervised fine-tuning, and reinforcement learning. Experiments show that DiaScriber achieves superior performance over comparison methods across extensive multi-speaker scenario test sets and demonstrates outstanding generalization ability in unseen multi-speaker scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。