7000万参数模型实现高效高质语音合成,支持自然表达与零样本克隆。
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model
- 分层解码架构以12赫兹速率生成语音,兼顾长文本建模与还原质量。
- 仅7000万参数即达到大模型同等音质,减少重复输出并提升稳定性。
- 适合需要轻量级、高鲁棒性语音合成的部署场景,如移动端应用。
基于编码器-解码器的语音合成模型在零样本语音克隆方面表现出色。然而,它们在处理更具表现力的参考语音或复杂文本输入时仍存在挑战。我们提出MARS6,一种高效的编码器-解码器变换器模型,支持快速、富有表现力的语音合成。MARS6基于近期语音语言建模的进展,采用分层解码结构,新语音标记以仅12赫兹的速率进行处理,从而实现对长文本的有效建模,同时保持高质量重建。我们结合多种最新训练与推理技术,降低重复生成,提升输出稳定性和音质。这使得7000万参数的MARS6在客观与主观评估中均达到远超其规模的性能,媲美数十倍大的模型。我们在语音合成质量及参考说话人克隆能力上进行了对比验证。
原文摘要 · Abstract (English)
Codec-based text-to-speech (TTS) models have shown impressive quality with zero-shot voice cloning abilities. However, they often struggle with more expressive references or complex text inputs. We present MARS6, a robust encoder-decoder transformer for rapid, expressive TTS. MARS6 is built on recent improvements in spoken language modelling. Utilizing a hierarchical setup for its decoder, new speech tokens are processed at a rate of only 12 Hz, enabling efficient modelling of long-form text while retaining reconstruction quality. We combine several recent training and inference techniques to reduce repetitive generation and improve output stability and quality. This enables the 70M-parameter MARS6 to achieve similar performance to models many times larger. We show this in objective and subjective evaluations, comparing TTS output quality and reference speaker cloning ability. Project page: https://camb-ai.github.io/mars6-turbo/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。