arXiv:2411.07691cs.AI2024-11综述被引 4

系统梳理预训练模型的安全与隐私风险,提出三类攻防分类框架。

New Emerged Security and Privacy of Pre-trained Model: a Survey and Outlook

  • 按输入和权重可访问性将攻击分为无改变、输入变更、模型变更三类。
  • 揭示预训练模型在隐私泄露和有害输出方面的独特安全挑战。
  • 适合关注大模型安全的科研人员和从业者参考。

得益于数据爆炸式增长和算力进步,预训练模型在自然语言处理、计算机视觉等任务中表现出卓越性能。然而,其广泛应用也引发了新的安全与隐私问题,如隐私信息泄露、生成有害内容等,严重削弱用户信任。随着模型性能持续提升,相关风险日益凸显。当前研究缺乏对新兴攻击与防御方法的清晰分类体系,制约了对该领域的深入理解。为此,本文系统综述预训练模型的安全风险,基于不同场景下输入与权重的可访问性,提出一套涵盖无改变、输入变更、模型变更三类的攻防分类框架。通过该框架,归纳总结现有安全问题的特征与本质,分析各类方法的优势与局限,并展望未来潜在的研究方向。

原文摘要 · Abstract (English)

Thanks to the explosive growth of data and the development of computational resources, it is possible to build pre-trained models that can achieve outstanding performance on various tasks, such as neural language processing, computer vision, and more. Despite their powerful capabilities, pre-trained models have also sparked attention to the emerging security challenges associated with their real-world applications. Security and privacy issues, such as leaking privacy information and generating harmful responses, have seriously undermined users' confidence in these powerful models. Concerns are growing as model performance improves dramatically. Researchers are eager to explore the unique security and privacy issues that have emerged, their distinguishing factors, and how to defend against them. However, the current literature lacks a clear taxonomy of emerging attacks and defenses for pre-trained models, which hinders a high-level and comprehensive understanding of these questions. To fill the gap, we conduct a systematical survey on the security risks of pre-trained models, proposing a taxonomy of attack and defense methods based on the accessibility of pre-trained models' input and weights in various security test scenarios. This taxonomy categorizes attacks and defenses into No-Change, Input-Change, and Model-Change approaches. With the taxonomy analysis, we capture the unique security and privacy issues of pre-trained models, categorizing and summarizing existing security issues based on their characteristics. In addition, we offer a timely and comprehensive review of each category's strengths and limitations. Our survey concludes by highlighting potential new research opportunities in the security and privacy of pre-trained models.

预训练模型安全隐私攻防分类综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。