arXiv:2604.26281eess.AScs.LG2026-04中稿 · Interspeech 2026

用扩散模型实现语音匿名化中韵律的可调控制,兼顾隐私与语音质量。

DiffAnon: Diffusion-based Prosody Control for Voice Anonymization

论文配图:DiffAnon: Diffusion-based Prosody Control for Voice Anonymization
图 1 · 摘自论文原文
  • 基于扩散模型与无分类器引导,实现推理时对韵律保留程度的连续控制。
  • 在单一模型中支持匿名强度与韵律保真度的平滑插值,实验验证了有效权衡。
  • 适合需要灵活调节隐私与语音自然度的应用场景,如语音保护与内容生成。

语音匿名化中是否保留韵律是一个核心问题。韵律传递语义与情感,但又与说话人身份紧密关联。现有方法要么为隐私牺牲韵律,要么缺乏对效用-隐私权衡的系统性控制机制,仅在固定设计点运行。本文提出 DiffAnon,一种基于扩散模型并结合无分类器引导(CFG)的匿名化方法,可在推理阶段实现对韵律保留的显式、连续控制。DiffAnon 在 RVQ 编码器的语义嵌入基础上精炼声学细节,使匿名化强度与韵律保真度之间的转换在单个模型内可平滑插值。据我们所知,这是首个提供结构化、可插值推理时韵律控制的语音匿名化框架。实验表明其具有结构化的权衡行为,在保持强语音质量的同时,于可控操作点上实现了具有竞争力的隐私保护效果。

原文摘要 · Abstract (English)

To preserve or not to preserve prosody is a central question in voice anonymization. Prosody conveys meaning and affect, yet is tightly coupled with speaker identity. Existing methods either discard prosody for privacy or lack a principled mechanism to control the utility-privacy trade-off, operating at fixed design points. We propose DiffAnon, a diffusion-based anonymization method with classifier-free guidance (CFG) that provides explicit, continuous inference-time control over prosody preservation. DiffAnon refines acoustic detail over semantic embeddings of an RVQ codec, enabling smooth interpolation between anonymization strength and prosodic fidelity within a single model. To the best of our knowledge, it is the first voice anonymization framework to provide structured, interpolatable inference-time prosody control. Experiments demonstrate structured trade-off behavior, achieving strong utility while maintaining competitive privacy across controllable operating points.

语音匿名扩散模型韵律控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。