arXiv:2510.25566eess.AScs.LG2025-10

PitchFlower让音频编码可精准控制音高,音质还更高。

PitchFlower: A flow-based neural audio codec with pitch controllability

  • 用随机平移音高轮廓训练,强制音高与其他特征解耦。
  • 音高控制精度超越WORLD,音质显著优于SiFiGAN。
  • 框架简单通用,适合拓展控制语调、语速等语音属性。

我们提出 PitchFlower,一种基于流的神经音频编解码器,具备显式的音高可控性。通过在训练中对基频(F0)轮廓进行平滑和随机偏移,并以真实 F0 作为条件输入,实现特征解耦。采用向量量化瓶颈阻止音高信息恢复,再由基于流的解码器生成高质量音频。实验表明,PitchFlower 在音高控制精度上超过 WORLD,且音质显著提升;在可控性上优于 SiFiGAN,同时保持相近音质。该框架为分离其他语音属性(如语调、语速)提供了简单且可扩展的路径。

原文摘要 · Abstract (English)

We present PitchFlower, a flow-based neural audio codec with explicit pitch controllability. Our approach enforces disentanglement through a simple perturbation: during training, F0 contours are flattened and randomly shifted, while the true F0 is provided as conditioning. A vector-quantization bottleneck prevents pitch recovery, and a flow-based decoder generates high quality audio. Experiments show that PitchFlower achieves more accurate pitch control than WORLD at much higher audio quality, and outperforms SiFiGAN in controllability while maintaining comparable quality. Beyond pitch, this framework provides a simple and extensible path toward disentangling other speech attributes.

音频编码音高控制流模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。