研究语言模型与视觉语言模型的后门攻击,揭示安全漏洞并提出防御思路。
Backdoor Learning in Language Models and Vision-Language Models

- 分析语言模型与多模态模型中的后门攻击机制。
- 发现攻击可隐蔽植入且在特定触发词下生效。
- 适合关注AI安全、模型可信性的研究人员阅读。
深度学习的进展显著提升了自然语言处理(NLP)和视觉语言模型(VLMs)的能力,但随之而来的安全漏洞也日益突出,尤其是后门攻击带来的严重威胁。本文从可信人工智能与高效多模态表征学习两个关键维度展开研究:(1) 安全性方面,系统分析、检测并设计针对NLP和VLMs的后门攻击;(2) 效率方面,开发适用于临床与医学影像的先进多模态表征方法。研究揭示了后门攻击在模型中的隐蔽植入路径,并验证其在特定触发词下的有效激活,为构建更安全可靠的智能系统提供理论支持。
原文摘要 · Abstract (English)
Recent advances in deep learning have significantly enhanced the capabilities of Natural Language Processing (NLP) and Vision-Language Models (VLMs). However, these advancements come with increased vulnerabilities, notably through backdoor attacks that pose severe security threats. This thesis addresses two critical dimensions of Trustworthy AI and Efficient Multimodal Representation Learning: (1) security through analyzing, detecting, and designing backdoor attacks in NLP and VLMs, and (2) efficiency through advanced multimodal representation methods tailored for clinical and medical imaging applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。