让语音生成更贴近人耳感知,提升小数据和少步数下的质量
LP-CFM: Perceptual Invariance-Aware Conditional Flow Matching for Speech Modeling
- 用线性投影建模语音的感知等价变体,使生成更符合人耳认知
- 在低资源和少步采样下,语音质量显著优于传统方法
- 适合语音合成、神经声码器等需感知鲁棒性的场景
本文提出一种新视角的语音建模方法,引入幅度缩放和时移等感知不变性。传统生成模型将每个样本视为目标分布的固定代表,但从生成角度看,这些样本只是真实语音分布中众多感知等价变体之一。为此,我们提出线性投影条件流匹配(LP-CFM),将目标建模为沿感知等价变体对齐的拉长高斯分布,并引入向量校准采样(VCS)以保持采样过程与线性投影路径一致。在不同模型规模、数据量和采样步数的神经声码器实验中,该方法始终优于传统最优传输条件流匹配(CFM),尤其在低资源和少步场景下提升显著。结果表明,LP-CFM与VCS能实现更鲁棒、更符合感知的语音生成建模。
原文摘要 · Abstract (English)
The goal of this paper is to provide a new perspective on speech modeling by incorporating perceptual invariances such as amplitude scaling and temporal shifts. Conventional generative formulations often treat each dataset sample as a fixed representative of the target distribution. From a generative standpoint, however, such samples are only one among many perceptually equivalent variants within the true speech distribution. To address this, we propose Linear Projection Conditional Flow Matching (LP-CFM), which models targets as projection-aligned elongated Gaussians along perceptually equivalent variants. We further introduce Vector Calibrated Sampling (VCS) to keep the sampling process aligned with the line-projection path. In neural vocoding experiments across model sizes, data scales, and sampling steps, the proposed approach consistently improves over the conventional optimal transport CFM, with particularly strong gains in low-resource and few-step scenarios. These results highlight the potential of LP-CFM and VCS to provide more robust and perceptually grounded generative modeling of speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。