arXiv:2510.23320eess.AScs.CL2025-10中稿 · TSD 2026被引 1

用文学文本生成对话数据,提升语音识别与说话人分离效果

LibriConvo: Simulating Conversations from Read Literature for ASR and Diarization

  • 基于文学文本构建合成对话,优化时序和语义连续性
  • 对话数据达240.1小时,说话人分离错误率降至11.1%
  • 适合研究合成对话数据的语音系统评估者使用

我们提出LibriConvo,一个用于说话人分离和自动语音识别(ASR)的合成对话语音语料库,基于先前提出的说话人感知模拟对话(SASC)框架构建。为提升下游任务适配性,通过外部语音活动检测估算通话时序统计,压缩长停顿,按书籍分组LibriTTS语句以增强局部语义连贯性,并采用空间合理性启发式选择房间冲激响应。最终语料包含240.1小时音频,涵盖1,496段对话、830名说话人,划分为互不重叠的训练、验证和测试集。报告了说话人分离与ASR基线结果:在测试集上,Sortformer在说话人分离上优于pyannote管道(11.1% vs. 24.4% DER);ASR方面,经序列输出训练微调的Fast Conformer-CTC XLarge模型达到7.29% WER和6.97% cpWER,优于零样本Whisper-large-v3。这些结果表明LibriConvo是研究合成对话语音与评估多说话人语音处理系统的实用基准。

原文摘要 · Abstract (English)

We introduce LibriConvo, a synthetic conversational speech corpus for speaker diarization and automatic speech recognition (ASR), built by instantiating the previously proposed Speaker-Aware Simulated Conversation (SASC) framework in a dataset and benchmarking setting. The main contribution of this paper is a corpus construction pipeline and benchmark derived from that framework. To make the data more suitable for downstream ASR and diarization, conversational timing statistics are estimated from English CallHome using external voice activity detection, long pauses are compressed, LibriTTS utterances are grouped by book to improve local semantic continuity, and room impulse responses are selected with a spatial-plausibility heuristic. The resulting corpus contains 240.1 hours of audio across 1,496 dialogues involving 830 speakers, partitioned into speaker-disjoint train, validation, and test splits. We report baseline results for both diarization and ASR. On the test split, Sortformer outperforms the pyannote pipeline in diarization (11.1\% vs.~24.4\% DER). For ASR, a Fast Conformer-CTC XLarge model fine-tuned with Serialized Output Training achieves 7.29\% WER and 6.97\% cpWER, outperforming zero-shot Whisper-large-v3. These results position LibriConvo as a practical benchmark for studying synthetic conversational speech and for evaluating multi-speaker speech processing systems.

语音识别说话人分离合成数据对话生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。