arXiv:2604.03074eess.AScs.CL2026-04被引 2

让语音模型像人一样多轮推理,精准识别多人对话中的说话人和时间点。

Speaker-Reasoner: Scaling Interaction Turns and Reasoning Patterns for Timestamped Speaker-Attributed ASR

  • 通过多轮迭代分析音频结构,自主预测说话时段边界。
  • 在阿里会议和AISHELL-4数据集上显著提升重叠语音与快速换人场景的识别准确率。
  • 适合需要高精度多人对话转写的场景,如会议记录、司法录音分析。

多说话人对话的转录与理解需要语音识别、说话人归属和时间戳定位。尽管语音大模型在单说话人任务中表现优异,但在多说话人场景下仍面临重叠语音、回应语、快速换人及上下文窗口限制等挑战。我们提出Speaker-Reasoner,一种具备代理式多轮时序推理能力的端到端语音大模型。该模型不采用单次推理,而是迭代分析全局音频结构,自主预测时间边界,并对细粒度片段进行联合建模,同时输出说话人身份、性别、时间戳和转写文本。引入说话人感知缓存机制,可处理超出训练上下文窗口的长音频。通过三阶段渐进式训练策略,Speaker-Reasoner在AliMeeting和AISHELL-4数据集上持续优于强基线,尤其在重叠语音和复杂换人场景中表现突出。

原文摘要 · Abstract (English)

Transcribing and understanding multi-speaker conversations requires speech recognition, speaker attribution, and timestamp localization. While speech LLMs excel at single-speaker tasks, multi-speaker scenarios remain challenging due to overlapping speech, backchannels, rapid turn-taking, and context window constraints. We propose Speaker-Reasoner, an end-to-end Speech LLM with agentic multi-turn temporal reasoning. Instead of single-pass inference, the model iteratively analyzes global audio structure, autonomously predicts temporal boundaries, and performs fine-grained segment analysis, jointly modeling speaker identity, gender, timestamps, and transcription. A speaker-aware cache further extends processing to audio exceeding the training context window. Trained with a three-stage progressive strategy, Speaker-Reasoner achieves consistent improvements over strong baselines on AliMeeting and AISHELL-4 datasets, particularly in handling overlapping speech and complex turn-taking.

语音识别多说话人时序推理对话分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。