arXiv:2606.09048eess.AScs.AI2026-06被引 1

直接生成语音波形,无需中间表示,提升音质与效率。

BareWave: Waveform-Native Flow-Matching Text-to-Speech

论文配图:BareWave: Waveform-Native Flow-Matching Text-to-Speech
图 1 · 摘自论文原文
  • 训练时对齐表征、分阶段降噪、感知对齐,提升优化效率。
  • 零样本克隆下语音自然度、发音清晰度和说话人相似性均达高水平。
  • 适合追求端到端语音生成性能的开发者与研究者。

去除中间表示和独立训练的解码阶段已成为生成建模的重要方向。然而,在文本转语音任务中,高质量系统仍普遍依赖中间声学表示后再进行波形合成。本文提出 BareWave,一种完全波形原生的流匹配语音合成框架,实现从文本到波形的直接生成。该设定带来三大挑战:原始波形建模缺乏强预训练表征支撑,不同训练阶段需适配不同噪声调度策略,且数据空间感知目标无法自动共享速度空间流目标的时间结构。为此,我们设计了一种直接文本到波形训练框架,结合训练时表征对齐、分阶段噪声调度和速度感知感知对齐(VAPA),同时保持测试时单一波形原生推理路径,无需预训练组件。在零样本语音克隆实验中,该方法在不依赖预训练模型的前提下,实现了高可懂度、强说话人相似性和自然语音质量,验证了波形原生流匹配语音合成的可行性与实用性。项目页面及音频演示见 https://barewave.github.io/。

原文摘要 · Abstract (English)

Removing intermediate representations and separately trained decoding stages has become an important direction in generative modeling. In text-to-speech, however, high-quality systems are still commonly built through an intermediate acoustic representation before waveform synthesis. In this work, we present BareWave, a fully waveform-native framework for direct text-to-wave generation in flow-matching TTS. We consider this setting to raise three training challenges: raw-waveform modeling lacks a strong pretrained representational scaffold, different stages of training benefit from different noise schedules, and data-space perceptual objectives do not automatically share the temporal structure of the velocity-space flow objective. As a result, direct waveform training is hard to optimize efficiently, hard to push toward a strong final operating point with a fixed recipe, and hard to integrate effective perceptual refinement. Guided by this view, we develop a direct text-to-wave training framework that combines training-time representation alignment, staged noise scheduling, and velocity-aware perceptual alignment (VAPA), while preserving a single waveform-native inference path without pretrained components at test time. Experiments on zero-shot voice cloning show that strong intelligibility, speaker similarity, and naturalness can be achieved under a fully waveform-native inference path, supporting waveform-native flow-matching TTS as a practical direction. Project page with audio demos is available at https://barewave.github.io/.

语音合成流匹配端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。