arXiv:2409.00946cs.SDcs.AI2024-09中稿 · the WI-IAT'24被引 3

用大模型生成多样真实对话音频,提升语音识别与分类模型训练效果

A Framework for Synthetic Audio Conversations Generation using Large Language Models

  • 基于大模型生成多角色对话文本,再通过语音合成转为音频
  • 生成数据在多样性与逼真度上表现优异,显著提升模型性能
  • 适合语音识别、音频分类等任务的模型训练与评估

本文提出 ConversaSynth 框架,利用大语言模型(LLMs)生成具有多种角色设定的合成对话音频。该框架首先生成跨多个主题的多样化且连贯的文本对话,随后通过文本转语音(TTS)系统转换为音频。实验表明,ConversaSynth 能有效生成高质量的合成音频数据集,显著提升音频标记、音频分类及多说话人语音识别模型的训练与评估效果。结果表明,该框架生成的数据具备显著的多样性与真实性,适用于构建鲁棒、可适应的语音人工智能系统。

原文摘要 · Abstract (English)

In this paper, we introduce ConversaSynth, a framework designed to generate synthetic conversation audio using large language models (LLMs) with multiple persona settings. The framework first creates diverse and coherent text-based dialogues across various topics, which are then converted into audio using text-to-speech (TTS) systems. Our experiments demonstrate that ConversaSynth effectively generates highquality synthetic audio datasets, which can significantly enhance the training and evaluation of models for audio tagging, audio classification, and multi-speaker speech recognition. The results indicate that the synthetic datasets generated by ConversaSynth exhibit substantial diversity and realism, making them suitable for developing robust, adaptable audio-based AI systems.

语音合成大模型对话生成数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。