arXiv:2601.10345cs.SD2026-01

用轻量扩散模型修复音高偏移造成的歌声失真,效果优于传统方法。

Self-supervised restoration of singing voice degraded by pitch shifting using shallow diffusion

  • 将音高偏移视为修复问题,用声学特征驱动轻量扩散模型重建歌声。
  • 在自监督训练下,相比经典方法显著降低音高偏移带来的失真。
  • 适合需要高质量音高变换的音乐制作场景,尤其关注自然度的用户。

音高偏移是歌唱音频处理中的关键技术,但传统信号处理方法存在形式共振偏移和机械感音色等固有缺陷,且在大跨度偏移时更为严重。本文提出将音高偏移重构为一种恢复问题:给定一个已被音高偏移(并引入伪影)的音频片段,目标是恢复出自然发声表现,同时保留旋律与节奏。具体采用一种轻量级、基于梅尔频谱空间的扩散模型,由帧级声学特征(如基频f0、音量、内容特征)驱动。通过自监督方式构建训练对:对真实音频施加音高偏移并反向还原,以模拟真实伪影并保留真实标签。在精心筛选的歌唱数据集上,所提方法在统计指标和成对声学测量中均显著优于代表性经典基线。结果表明,基于恢复的音高偏移方法有望成为抗伪影语音转调工作流的可行方案。

原文摘要 · Abstract (English)

Pitch shifting has been an essential feature in singing voice production. However, conventional signal processing approaches exhibit well known trade offs such as formant shifts and robotic coloration that becomes more severe at larger transposition jumps. This paper targets high quality pitch shifting for singing by reframing it as a restoration problem: given an audio track that has been pitch shifted (and thus contaminated by artifacts), we recover a natural sounding performance while preserving its melody and timing. Specifically, we use a lightweight, mel space diffusion model driven by frame level acoustic features such as f0, volume, and content features. We construct training pairs in a self supervised manner by applying pitch shifts and reversing them to simulate realistic artifacts while retaining ground truth. On a curated singing set, the proposed approach substantially reduces pitch shift artifacts compared to representative classical baselines, as measured by both statistical metrics and pairwise acoustic measures. The results suggest that restoration based pitch shifting could be a viable approach towards artifact resistant transposition in vocal production workflows.

音高偏移扩散模型歌声修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。