arXiv:2604.09246cs.SDcs.AI2026-04中稿 · CHI workshop

改进语音匿名化中异常语音的音质,减少刺耳杂音。

DDSP-QbE++: Improving Speech Quality for Speech Anonymisation for Atypical Speech

论文配图:DDSP-QbE++: Improving Speech Quality for Speech Anonymisation for Atypical Speech
图 1 · 摘自论文原文
  • 用声门检测控制谐波激励,未发声区改用滤波噪声。
  • 引入多项式带限步进修正,消除相位跳变引起的失真。
  • 无需额外参数,可直接融入现有训练流程,适合语音匿名化研究者。

基于可微分数字信号处理(DDSP)的语音转换依赖减法合成,通过学习频谱包络对周期性激励信号进行整形以重建目标语音。在DDSP-QbE中,激励信号通过相位累加生成锯齿波,其突变不连续性会引入混叠伪影,表现为听觉上的刺耳感和频谱失真,尤其在较高基频时更为明显。本文提出两项针对激励阶段的改进:首先,引入显式声门检测机制,在非发声区域抑制周期性成分,并以滤波噪声替代,避免最易被感知的混叠谐波;其次,采用多项式带限步进(PolyBLEP)校正相位累加振荡器,在每次相位重置处用平滑多项式残差替代硬跳变,有效消除产生混叠的成分,且无需过采样或频谱截断。两项改进协同作用,实现更平滑的谐波滚降、更低的高频伪影及更高的主观自然度(经MOS评分验证)。所提方法轻量、可微分,无缝集成于原有DDSP-QbE训练流程,无新增可学习参数。

原文摘要 · Abstract (English)

Differentiable Digital Signal Processing (DDSP) pipelines for voice conversion rely on subtractive synthesis, where a periodic excitation signal is shaped by a learned spectral envelope to reconstruct the target voice. In DDSP-QbE, the excitation is generated via phase accumulation, producing a sawtooth-like waveform whose abrupt discontinuities introduce aliasing artefacts that manifest perceptually as buzziness and spectral distortion, particularly at higher fundamental frequencies. We propose two targeted improvements to the excitation stage of the DDSP-QbE subtractive synthesizer. First, we incorporate explicit voicing detection to gate the harmonic excitation, suppressing the periodic component in unvoiced regions and replacing it with filtered noise, thereby avoiding aliased harmonic content where it is most perceptually disruptive. Second, we apply Polynomial Band-Limited Step (PolyBLEP) correction to the phase-accumulated oscillator, substituting the hard waveform discontinuity at each phase wrap with a smooth polynomial residual that cancels alias-generating components without oversampling or spectral truncation. Together, these modifications yield a cleaner harmonic roll-off, reduced high-frequency artefacts, and improved perceptual naturalness, as measured by MOS. The proposed approach is lightweight, differentiable, and integrates seamlessly into the existing DDSP-QbE training pipeline with no additional learnable parameters.

语音匿名化音质提升信号处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。