arXiv:2501.15442cs.SDcs.AI2025-01综述被引 12

开源音频生成工具包v0.2发布,支持多语言语音合成等任务

Overview of the Amphion Toolkit (v0.2)

  • 提供多语言音频生成框架与数据处理流水线
  • 包含10万小时多语言开源数据集及新模型
  • 适合初学者快速上手语音与音乐生成研究

Amphion 是一个面向音频、音乐和语音生成的开源工具包,旨在降低该领域初学者的研究门槛。本文介绍2024年推出的Amphion v0.2版本,该版本包含10万小时的开源多语言数据集、可靠的预处理流程,以及文本到语音、音频编码、语音转换等任务的新模型。报告还提供了多个教程,指导用户使用新模型的功能与方法。

原文摘要 · Abstract (English)

Amphion is an open-source toolkit for Audio, Music, and Speech Generation, designed to lower the entry barrier for junior researchers and engineers in these fields. It provides a versatile framework that supports a variety of generation tasks and models. In this report, we introduce Amphion v0.2, the second major release developed in 2024. This release features a 100K-hour open-source multilingual dataset, a robust data preparation pipeline, and novel models for tasks such as text-to-speech, audio coding, and voice conversion. Furthermore, the report includes multiple tutorials that guide users through the functionalities and usage of the newly released models.

语音生成开源工具多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。