arXiv:2508.05554cs.SDcs.CL2025-08被引 7

金融领域多说话人语音转录数据集,支持精准说话人标注。

SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription

  • 构建金融会议录音的多说话人转录数据集,含完整文本与说话人信息。
  • 新增3780小时专业转录音频,提升端到端语音识别性能。
  • 适合研究金融语音识别、说话人分离及多任务建模的学者使用。

我们提出SPGISpeech 2.0,一个适用于金融领域说话人标注转录的数据集。该数据集在保持原始SPGISpeech核心特性的基础上,扩展了可适用建模任务的多样性:包含音频片段及其对应的完整格式化文本转录,可用于端到端自动语音识别(ASR)。SPGISpeech 2.0新增了3,780小时经过专业转录的财报电话会议音频,并为每个音频片段提供了通话和说话人信息,支持多说话人语音识别。通过在SPGISpeech 2.0上微调主流语音识别模型,验证了其在说话人标注语音识别任务中的有效性。数据集免费用于非商业用途,预计将推动语音识别技术发展并激发广泛研究应用。

原文摘要 · Abstract (English)

We introduce SPGISpeech 2.0, a dataset suitable for speaker-tagged transcription in the financial domain. SPGISpeech 2.0 improves the diversity of applicable modeling tasks while maintaining the core characteristic of the original SPGISpeech dataset: audio snippets and their corresponding fully formatted text transcriptions, usable for end-to-end automatic speech recognition (ASR). SPGISpeech 2.0 consists of 3,780 additional hours of professionally transcribed earnings calls. Furthermore, the dataset contains call and speaker information for each audio snippet facilitating multi-talker ASR. We validate the utility of SPGISpeech 2.0 through improvements in speaker-tagged ASR performance of popular speech recognition models after fine-tuning on SPGISpeech 2.0. Released free for non-commercial use, we expect SPGISpeech 2.0 to foster advancements in speech recognition technologies and inspire a wide range of research applications.

语音识别金融数据多说话人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。