用多模态大模型从音频推断隐私属性,揭示新型数据泄露风险。
The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents
- 构建混合智能体框架Gifts,结合语音-语言模型与大语言模型提升推理能力
- 在自建的AP²数据集上,该方法对敏感属性的推断准确率显著高于基线
- 适用于隐私安全研究者,尤其关注音频数据泄露防护的团队
本研究揭示了多模态大语言模型(MLLMs)的一项新型隐私风险:可从音频数据中推断敏感个人属性,我们称之为音频隐私属性画像。相比图像和文本,音频具有音调、语调等独特特征,更易用于精细画像,且可隐蔽采集。然而,现有研究面临两大挑战:缺乏带敏感属性标注的音频基准数据集,以及当前MLLM难以直接从音频中推断属性。为此,我们构建了AP²数据集,包含两个真实世界数据来源的子集,并带有敏感属性标签;提出Gifts框架,通过大语言模型(LLM)引导音频-语言模型(ALM)进行推理,并对其输出进行取证分析与整合,有效缓解了现有ALM在长上下文生成中的严重幻觉问题。实验表明,Gifts显著优于基线方法。最后,我们探索了模型级与数据级防御策略以缓解风险。本工作验证了基于MLLM的音频隐私攻击可行性,强调需建立强防御机制,并为后续研究提供数据与工具支持。
原文摘要 · Abstract (English)
Our research uncovers a novel privacy risk associated with multimodal large language models (MLLMs): the ability to infer sensitive personal attributes from audio data -- a technique we term audio private attribute profiling. This capability poses a significant threat, as audio can be covertly captured without direct interaction or visibility. Moreover, compared to images and text, audio carries unique characteristics, such as tone and pitch, which can be exploited for more detailed profiling. However, two key challenges exist in understanding MLLM-employed private attribute profiling from audio: (1) the lack of audio benchmark datasets with sensitive attribute annotations and (2) the limited ability of current MLLMs to infer such attributes directly from audio. To address these challenges, we introduce AP^2, an audio benchmark dataset that consists of two subsets collected and composed from real-world data, and both are annotated with sensitive attribute labels. Additionally, we propose Gifts, a hybrid multi-agent framework that leverages the complementary strengths of audio-language models (ALMs) and large language models (LLMs) to enhance inference capabilities. Gifts employs an LLM to guide the ALM in inferring sensitive attributes, then forensically analyzes and consolidates the ALM's inferences, overcoming severe hallucinations of existing ALMs in generating long-context responses. Our evaluations demonstrate that Gifts significantly outperforms baseline approaches in inferring sensitive attributes. Finally, we investigate model-level and data-level defense strategies to mitigate the risks of audio private attribute profiling. Our work validates the feasibility of audio-based privacy attacks using MLLMs, highlighting the need for robust defenses, and provides a dataset and framework to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。