arXiv:2605.05231eess.AScs.SD2026-05

用带说话人标签的提示词微调Whisper,实现高精度语音转写与说话人分离。

Prompting Whisper for Joint Speech Transcription and Diarization

  • 用带说话人标签的提示词微调Whisper,生成类似SOT格式的输出。
  • 相比原始Whisper,长音频中说话人标识更一致,转录更准确。
  • 发现提示词错误会传递,重叠语音的时间戳标注仍不准确。

作为MediSpeech项目的一部分,我们致力于开发一个实时转写并分离医生与患者对话的系统。本研究(进行中)探索如何高效结合Whisper与说话人分离(SD)。在尝试用含说话人标签的文本提示Whisper后,发现其能以较高准确率插入标签。随后通过使用带说话人标签的提示词对Whisper进行微调,生成类似序列化输出训练(SOT)格式的转写结果。微调后的Whisper在长音频中表现出更一致的说话人标识和更精确的逐字转写。但研究也揭示新挑战:由于提示词错误传播,Whisper的SD性能下降,且重叠语音的时间戳标注不准确。

原文摘要 · Abstract (English)

As part of the MediSpeech project, we aim to develop a system that transcribes and diarizes Dutch conversations between doctors and patients in real-time. In this research (in-progress) we explore ways of efficiently combining Whisper with speaker diarization (SD). After trying to prompt Whisper with text that contains speaker labels, we observed that it is able to insert labels into the transcription with promising accuracy. We continued this line of research by fine-tuning Whisper with speaker-labelled prompts to generate transcriptions in a format similar to that of Serialized Output Training (SOT). Fine-tuning Whisper yielded more consistent speaker IDs across the chunks of long-form audio and improved verbatim transcription. The study uncovered new challenges as Whisper's SD performance suffers because of mistakes that get propagated through prompts and inaccurate timestamps assigned to overlapping speech.

语音转写说话人分离WhisperSOT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。