arXiv:2506.05796eess.AS2025-06被引 11

用大模型融合说话人识别,让语音转写更准且带时间信息。

Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models

  • 用大模型同时处理说话人身份和语义,实现端到端转写。
  • 在高重叠会议场景中表现优异,多语言对话准确率显著提升。
  • 适合需要精准说话人分离的会议记录、对话分析等场景。

多说话人自动语音识别(MS-ASR)在处理重叠语音时面临重大挑战,这对会议转录和对话分析等应用至关重要。尽管串行输出训练(SOT)方法是常见方案,但常丢弃绝对时间信息,限制了其在时间敏感场景的应用。借助大语言模型(LLM)在对话音频处理中的最新进展,我们提出一种新型的说话人感知多说话人语音识别系统,将说话人聚类与基于大模型的转写相结合。该框架结合结构化聚类输入、帧级说话人嵌入和语义嵌入,使大模型能够生成段落级转写结果。实验表明,该系统在多语言双人对话中表现稳健,并在复杂高重叠多说话人会议场景中表现出色。本工作展示了大模型作为统一后端,在联合说话人感知分割与转写方面的潜力。

原文摘要 · Abstract (English)

Multi-speaker automatic speech recognition (MS-ASR) faces significant challenges in transcribing overlapped speech, a task critical for applications like meeting transcription and conversational analysis. While serialized output training (SOT)-style methods serve as common solutions, they often discard absolute timing information, limiting their utility in time-sensitive scenarios. Leveraging recent advances in large language models (LLMs) for conversational audio processing, we propose a novel diarization-aware multi-speaker ASR system that integrates speaker diarization with LLM-based transcription. Our framework processes structured diarization inputs alongside frame-level speaker and semantic embeddings, enabling the LLM to generate segment-level transcriptions. Experiments demonstrate that the system achieves robust performance in multilingual dyadic conversations and excels in complex, high-overlap multi-speaker meeting scenarios. This work highlights the potential of LLMs as unified back-ends for joint speaker-aware segmentation and transcription.

语音识别大模型说话人分离会议转录

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。