用音调和节奏调整保护语音情感隐私,既安全又不影响使用。
Exploring Audio Editing Features as User-Centric Privacy Defenses Against Large Language Model(LLM) Based Emotion Inference Attacks
- 利用音调与节奏修改实现用户主导的隐私防护
- 在三个数据集上有效混淆情感信息,抵御多种攻击
- 适合注重隐私的语音应用开发者和普通用户
语音助手、视频会议平台和可穿戴设备等语音技术的普及,引发了敏感情感信息被音频数据推断的隐私担忧。现有隐私保护方法常牺牲可用性与安全性,难以实际应用。本文提出一种用户为中心的新方法,利用安卓和iOS平台上常见的音调与节奏编辑功能,在不降低可用性的前提下保护情感隐私。通过分析主流音频编辑应用,我们确认这些功能广泛可用且易于操作。针对来自深度神经网络、大语言模型及可逆性测试等多源对抗攻击的威胁模型,我们在三个数据集上进行了严格评估,结果表明音调与节奏调整能有效混淆情感信息。此外,论文还探讨了轻量级、设备端部署的设计原则,以确保跨平台、跨设备的广泛应用。
原文摘要 · Abstract (English)
The rapid proliferation of speech-enabled technologies, including virtual assistants, video conferencing platforms, and wearable devices, has raised significant privacy concerns, particularly regarding the inference of sensitive emotional information from audio data. Existing privacy-preserving methods often compromise usability and security, limiting their adoption in practical scenarios. This paper introduces a novel, user-centric approach that leverages familiar audio editing techniques, specifically pitch and tempo manipulation, to protect emotional privacy without sacrificing usability. By analyzing popular audio editing applications on Android and iOS platforms, we identified these features as both widely available and usable. We rigorously evaluated their effectiveness against a threat model, considering adversarial attacks from diverse sources, including Deep Neural Networks (DNNs), Large Language Models (LLMs), and and reversibility testing. Our experiments, conducted on three distinct datasets, demonstrate that pitch and tempo manipulation effectively obfuscates emotional data. Additionally, we explore the design principles for lightweight, on-device implementation to ensure broad applicability across various devices and platforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。