用大模型统一分析AI全生命周期隐私风险,兼顾传统与新型攻击。
PriMod4AI: Lifecycle-Aware Privacy Threat Modeling for AI Systems using LLM
- 结合LINDDUN与模型攻击知识库,构建混合威胁建模框架。
- 在两个AI系统上识别出传统隐私威胁及成员推理等新型攻击。
- 适合安全评估者、合规人员,提升隐私风险分析的全面性。
人工智能系统在其生命周期中引入复杂的隐私风险,尤其在处理敏感或高维数据时。除了LINDDUN框架定义的七类传统隐私威胁外,AI系统还面临成员推理、模型反演等以模型为中心的隐私攻击,而这些未被LINDDUN涵盖。为同时应对经典威胁与新兴AI驱动攻击,PriMod4AI提出一种混合隐私威胁建模方法,整合两个结构化知识源:代表既有分类体系的LINDDUN知识库,以及捕捉非LINDDUN威胁的模型中心隐私攻击知识库。这些知识库被嵌入向量数据库用于语义检索,并与来自数据流图的系统级元数据结合。通过检索增强与数据流特定提示生成,引导大语言模型(LLMs)在各生命周期阶段识别、解释并分类隐私威胁。该框架生成有依据且基于分类体系的威胁评估,融合经典与AI驱动视角。对两个AI系统的评估表明,PriMod4AI不仅覆盖了广泛的经典隐私类别,还额外识别出模型中心隐私威胁。框架在不同LLM间输出一致,体现于观测范围内的共识评分。
原文摘要 · Abstract (English)
Artificial intelligence systems introduce complex privacy risks throughout their lifecycle, especially when processing sensitive or high-dimensional data. Beyond the seven traditional privacy threat categories defined by the LINDDUN framework, AI systems are also exposed to model-centric privacy attacks such as membership inference and model inversion, which LINDDUN does not cover. To address both classical LINDDUN threats and additional AI-driven privacy attacks, PriMod4AI introduces a hybrid privacy threat modeling approach that unifies two structured knowledge sources, a LINDDUN knowledge base representing the established taxonomy, and a model-centric privacy attack knowledge base capturing threats outside LINDDUN. These knowledge bases are embedded into a vector database for semantic retrieval and combined with system level metadata derived from Data Flow Diagram. PriMod4AI uses retrieval-augmented and Data Flow specific prompt generation to guide large language models (LLMs) in identifying, explaining, and categorizing privacy threats across lifecycle stages. The framework produces justified and taxonomy-grounded threat assessments that integrate both classical and AI-driven perspectives. Evaluation on two AI systems indicates that PriMod4AI provides broad coverage of classical privacy categories while additionally identifying model-centric privacy threats. The framework produces consistent, knowledge-grounded outputs across LLMs, as reflected in agreement scores in the observed range.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。