量化单对一语音转换中的说话人信息泄露,揭示隐私风险
Quantifying Source Speaker Leakage in One-to-One Voice Conversion
- 用平行语料库测试语音转换后说话人特征泄露
- 在白盒攻击下可准确推断源说话人身份,缩小嫌疑人范围
- 提醒语音合成服务商需保障说话人数据隐私
基于多口音平行语句语料库,我们开展案例研究,表明在一对一语音转换中可量化对源说话人身份的置信度。采用HiFi-GAN声码器进行语音转换后,我们评估了多种说话人特征的信息泄露程度。在假设最坏情况的白盒攻击场景下,我们能够有效推断源说话人身份,并缩小可能的候选者范围。该结果强化了语音合成服务提供方在数据隐私方面的监管责任与道德义务。
原文摘要 · Abstract (English)
Using a multi-accented corpus of parallel utterances for use with commercial speech devices, we present a case study to show that it is possible to quantify a degree of confidence about a source speaker's identity in the case of one-to-one voice conversion. Following voice conversion using a HiFi-GAN vocoder, we compare information leakage for a range speaker characteristics; assuming a "worst-case" white-box scenario, we quantify our confidence to perform inference and narrow the pool of likely source speakers, reinforcing the regulatory obligation and moral duty that providers of synthetic voices have to ensure the privacy of their speakers' data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。