arXiv:2603.24857cs.CRcs.AI2026-03综述被引 4

首次构建模型与数据双向交互的统一安全框架,厘清大模型时代各类攻击的内在关联。

AI Security in the Foundation Model Era: A Comprehensive Survey from a Unified Perspective

  • 提出四维闭环威胁分类,揭示数据与模型间的双向风险传导机制。
  • 明确四类攻击:数据窃取、模型污染、数据反演、模型盗取,覆盖全链条风险。
  • 适合研究大模型安全的学者和工程师,为系统性防御提供理论基石。

随着机器学习系统规模与功能不断扩展,安全威胁日益复杂,攻击与防御手段层出不穷。然而,现有研究多将各类威胁孤立看待,缺乏统一框架揭示其共性规律与相互依赖关系,阻碍了系统性理解与全面防御设计。关键在于,机器学习的两大基础资产——数据与模型——已不再独立,一方漏洞会直接危及另一方。现有框架未能阐明这些双向风险如何在机器学习全流程中传播。为此,本文提出一个统一的闭环威胁分类体系,通过四个方向轴显式建模模型与数据的交互关系。该框架涵盖四类安全威胁:(1) 数据→数据(D→D):包括数据解密攻击与水印移除攻击;(2) 数据→模型(D→M):包括投毒攻击、有害微调攻击与越狱攻击;(3) 模型→数据(M→D):包括模型反演、成员推断攻击与训练数据提取攻击;(4) 模型→模型(M→M):包括模型提取攻击。该统一框架揭示了各类安全威胁的深层关联,为构建可扩展、可迁移、跨模态的安全策略奠定基础,尤其适用于基础模型场景。

原文摘要 · Abstract (English)

As machine learning (ML) systems expand in both scale and functionality, the security landscape has become increasingly complex, with a proliferation of attacks and defenses. However, existing studies largely treat these threats in isolation, lacking a coherent framework to expose their shared principles and interdependencies. This fragmented view hinders systematic understanding and limits the design of comprehensive defenses. Crucially, the two foundational assets of ML -- \textbf{data} and \textbf{models} -- are no longer independent; vulnerabilities in one directly compromise the other. The absence of a holistic framework leaves open questions about how these bidirectional risks propagate across the ML pipeline. To address this critical gap, we propose a \emph{unified closed-loop threat taxonomy} that explicitly frames model-data interactions along four directional axes. Our framework offers a principled lens for analyzing and defending foundation models. The resulting four classes of security threats represent distinct but interrelated categories of attacks: (1) Data$\rightarrow$Data (D$\rightarrow$D): including \emph{data decryption attacks and watermark removal attacks}; (2) Data$\rightarrow$Model (D$\rightarrow$M): including \emph{poisoning, harmful fine-tuning attacks, and jailbreak attacks}; (3) Model$\rightarrow$Data (M$\rightarrow$D): including \emph{model inversion, membership inference attacks, and training data extraction attacks}; (4) Model$\rightarrow$Model (M$\rightarrow$M): including \emph{model extraction attacks}. Our unified framework elucidates the underlying connections among these security threats and establishes a foundation for developing scalable, transferable, and cross-modal security strategies, particularly within the landscape of foundation models.

AI安全大模型威胁框架数据隐私

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。