arXiv:2506.00267cs.CLcs.SD2025-06被引 2

构建100+小时自然对话数据集,填补高质量口语数据空白

CASPER: A Large Scale Spontaneous Speech Dataset

  • 设计新采集流程,引导真实自然对话
  • 发布超100小时自发口语数据,覆盖多样话题
  • 提供可复现框架,适合语音模型研究者使用

大型语言模型的成功推动了语音处理能力的发展,但高质量自发口语数据稀缺,现有数据集多为脚本化对话。为此,我们提出一种新型对话采集与录制流程,发布包含100+小时自发口语的语料库。该方法促进流畅自然的交流,鼓励多样化话题和互动。相比传统方式,能更真实地反映人类对话模式,提供可复现的数据收集框架。本文介绍数据集与方法,为解决自发口语数据不足问题奠定基础。未来计划持续扩展,形成面向研究社区的动态资源。

原文摘要 · Abstract (English)

The success of large language models has driven interest in developing similar speech processing capabilities. However, a key challenge is the scarcity of high-quality spontaneous speech data, as most existing datasets contain scripted dialogues. To address this, we present a novel pipeline for eliciting and recording natural dialogues and release our dataset with 100+ hours of spontaneous speech. Our approach fosters fluid, natural conversations while encouraging a diverse range of topics and interactive exchanges. Unlike traditional methods, it facilitates genuine interactions, providing a reproducible framework for future data collection. This paper introduces our dataset and methodology, laying the groundwork for addressing the shortage of spontaneous speech data. We plan to expand this dataset in future stages, offering a growing resource for the research community.

语音数据自发对话数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。