用扩散模型强化学习,实现语音去标识同时保留认知健康信息。
DDPO-VC: Speaker De-Identification via Diffusion Denoising Policy Optimization

- 基于扩散模型与强化学习,通过隐私与效用双教师奖励优化语音表示。
- 在两个阿尔茨海默病语音数据集上,隐私保护与认知任务性能均优于现有方法。
- 适合语音隐私保护与医疗语音分析场景,兼顾安全与实用价值。
语音去标识的核心挑战在于隐私与效用的平衡。许多效用变量(如说话人认知健康状态)与隐私变量(如说话人身份)存在相关性,违反了解耦方法依赖的独立性假设,导致私密信息泄露及下游任务有用信息丢失。为此,我们提出一种通用框架DDPO-VC,通过基于强化学习的扩散模型后训练实现语音去标识。该方法从关注隐私与关注效用的双教师中获取奖励信号,有效提升隐私保护能力与认知功能预测的可用性。在两个常用的阿尔茨海默病语音基准数据集上,其表现显著优于多种强基线去标识方法,兼具高隐私安全性与高下游任务效用。
原文摘要 · Abstract (English)
A key challenge of speaker de-identification is the balance between privacy and utility. Many utility variables, such as the cognitive health status of the speaker, are correlated with the privacy variable, such as the speaker identity, violating the independence assumption held by the disentanglement-based approaches, causing leakage of private information and the loss of useful information for downstream tasks. To tackle this challenge, we propose a general framework, DDPO-VC, for speaker de-identification through reinforcement learning-based post-training with diffusion models. Learning from reward signals combining knowledge from privacy-focused and utility-focused teachers, our method outperforms various strong \deid/ methods in both privacy preservation and cognitive utility on two commonly used dementia speech benchmarks. Please check out our code\footnote{\href{https://github.com/cactuswiththoughts/DDPO-VC}{https://github.com/cactuswiththoughts/DDPO-VC}} and demo\footnote{\href{https://cactuswiththoughts.github.io/SpeakerDeID-Demo/}{https://cactuswiththoughts.github.io/SpeakerDeID-Demo/}}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。