构建大规模多说话人多语言对话语音数据集,支持研究真实对话场景。
OleSpeech-IV: A Large-Scale Multispeaker and Multilingual Conversational Speech Dataset with Diverse Topics
- 从公开播客等真实对话中采集音频,人工标注说话人与转录文本。
- 提供带时间戳和置信度评分的精细标注,覆盖多种话题与语言。
- 开源英文阿拉伯语子集,适合语音识别与对话系统研究者使用。
OleSpeech-IV 是一个大规模多说话人、多语言的对话语音数据集,内容源自公开的英语播客、访谈节目、电话会议等真实对话。说话人名称、发言轮次和文本转录由人工标注,并通过专有流程精炼;额外信息如时间戳和置信度分数也由该流程生成。其中 'IV' 表示其为 Olewave 数据集系列中的第四级。此外,我们已开放一个子集 OleSpeech-IV-2025-EN-AR-100,供非商业研究使用。
原文摘要 · Abstract (English)
OleSpeech-IV dataset is a large-scale multispeaker and multilingual conversational speech dataset with diverse topics. The audio content comes from publicly-available English podcasts, talk shows, teleconferences, and other conversations. Speaker names, turns, and transcripts are human-sourced and refined by a proprietary pipeline, while additional information such as timestamps and confidence scores is derived from the pipeline. The IV denotes its position as Tier IV in the Olewave dataset series. In addition, we have open-sourced a subset, OleSpeech-IV-2025-EN-AR-100, for non-commercial research use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。