端到端建模谁在何时说了什么,提升多人语音识别与说话人分离精度。
TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding
- 分语义与说话人双流设计,通过序列输出训练学习发言切换规律。
- 在AMI和AliMeeting上DER优于Qwen-Omni和Gemini,重叠语音处理更优。
- 仅训练轻量投影模块,保持大模型能力同时降低计算开销。
我们提出TagSpeech,一种基于大语言模型的统一框架,利用时间锚点对齐实现端到端多说话人语音识别与说话人分离。该框架包含两个核心设计:(1) 通过序列输出训练(SOT)微调解耦的语义与说话人流,以学习发言轮换动态;(2) 交错式时间锚点机制,不仅支持细粒度时间戳预测,还作为语义理解与说话人追踪之间的同步信号。相较于以往侧重说话人标注语音识别或隐式分离的工作,TagSpeech解决了细粒度说话人-内容对齐难题,显式建模“谁在何时说了什么”。在AMI和AliMeeting数据集上的实验表明,该方法在端到端基线(包括Qwen-Omni和Gemini)基础上持续提升说话人分离错误率(DER),尤其在复杂语音重叠场景下表现更优。此外,TagSpeech采用参数高效训练范式,冻结大模型主干,仅训练轻量投影模块,实现高性能与低计算成本的平衡。
原文摘要 · Abstract (English)
We present TagSpeech, a unified LLM-based framework that utilizes Temporal Anchor Grounding for joint multi-speaker ASR and diarization. The framework is built on two key designs: (1) decoupled semantic and speaker streams fine-tuned via Serialized Output Training (SOT) to learn turn-taking dynamics; and (2) an interleaved time anchor mechanism that not only supports fine-grained timestamp prediction but also acts as a synchronization signal between semantic understanding and speaker tracking. Compared to previous works that primarily focus on speaker-attributed ASR or implicit diarization, TagSpeech addresses the challenge of fine-grained speaker-content alignment and explicitly models "who spoke what and when" in an end-to-end manner. Experiments on AMI and AliMeeting benchmarks demonstrate that our method achieves consistent improvements in Diarization Error Rate (DER) over strong end-to-end baselines, including Qwen-Omni and Gemini, particularly in handling complex speech overlaps. Moreover, TagSpeech employs a parameter-efficient training paradigm in which the LLM backbone is frozen and only lightweight projectors are trained, resulting in strong performance with low computational cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。