arXiv:2503.23848cs.CL2025-03被引 2

自动生成高质量语音对话数据,降低语音大模型研发成本。

SpeechDialogueFactory: Generating High-Quality Speech Dialogue Data to Accelerate Your Speech-LLM Development

  • 用多阶段流程合成带语气的自然语音对话
  • 生成质量接近真人录音,成本大幅降低
  • 开源工具支持中英文数据,适合语音大模型开发者

高质量语音对话数据对语音大模型开发至关重要,但现有获取方法存在明显局限:人工录制成本高且涉及隐私问题,而合成方法往往缺乏对话真实感。为此,我们提出 extsc{SpeechDialogueFactory},一个可直接投入使用的高效语音对话生成框架。该系统包含元数据生成、对话脚本设计、带副语言特征的语句模拟以及基于语音克隆的自然语音合成等完整流程,并提供交互式界面用于样本检查及高吞吐批量合成模式。评估表明,系统生成的对话质量可媲美真人录音,同时显著降低制作成本。我们已将该工作开源,配套提供中英文示例数据集,助力语音大模型研究与开发。

原文摘要 · Abstract (English)

High-quality speech dialogue datasets are crucial for Speech-LLM development, yet existing acquisition methods face significant limitations. Human recordings incur high costs and privacy concerns, while synthetic approaches often lack conversational authenticity. To address these challenges, we introduce \textsc{SpeechDialogueFactory}, a production-ready framework for generating natural speech dialogues efficiently. Our solution employs a comprehensive pipeline including metadata generation, dialogue scripting, paralinguistic-enriched utterance simulation, and natural speech synthesis with voice cloning. Additionally, the system provides an interactive UI for detailed sample inspection and a high-throughput batch synthesis mode. Evaluations show that dialogues generated by our system achieve a quality comparable to human recordings while significantly reducing production costs. We release our work as an open-source toolkit, alongside example datasets available in English and Chinese, empowering researchers and developers in Speech-LLM research and development.

语音生成数据合成语音大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。