提出双触发后门攻击,无需优化文本属性即可在图模型中隐蔽植入恶意行为。
Stealthy Dual-Trigger Backdoors: Attacking Prompt Tuning in LM-Empowered Graph Foundation Models
- 设计文本与结构双触发机制,利用预存文本池实现无须优化的攻击
- 在不可见属性场景下仍保持高成功率,且不损害正常准确率
- 揭示语言模型增强图模型的隐蔽安全风险,适合关注AI安全的研究者
图基础模型(GFMs)尤其是融合语言模型(LMs)的版本,在文本属性图(TAGs)上表现出卓越性能。然而,相较于传统图神经网络,这类模型在未受保护的提示调优阶段引入了独特安全漏洞,当前研究尚未充分关注。实证发现,在无法访问属性的受限标签系统中,传统图后门攻击性能显著下降,尤其当触发节点属性无法显式优化时。为此,本文提出一种新型双触发后门攻击框架,同时作用于文本层与结构层,通过战略性使用预设文本池,在无需显式优化触发节点文本属性的情况下实现有效攻击。大量实验表明,该攻击在保持优异纯净准确率的同时,仍能实现极高的攻击成功率,包括高度隐蔽的单触发节点场景。本工作揭示了部署于网络的基于语言模型的图基础模型中的关键后门风险,有助于推动开源平台在基础模型时代建立更健壮的监督机制。
原文摘要 · Abstract (English)
The emergence of graph foundation models (GFMs), particularly those incorporating language models (LMs), has revolutionized graph learning and demonstrated remarkable performance on text-attributed graphs (TAGs). However, compared to traditional GNNs, these LM-empowered GFMs introduce unique security vulnerabilities during the unsecured prompt tuning phase that remain understudied in current research. Through empirical investigation, we reveal a significant performance degradation in traditional graph backdoor attacks when operating in attribute-inaccessible constrained TAG systems without explicit trigger node attribute optimization. To address this, we propose a novel dual-trigger backdoor attack framework that operates at both text-level and struct-level, enabling effective attacks without explicit optimization of trigger node text attributes through the strategic utilization of a pre-established text pool. Extensive experimental evaluations demonstrate that our attack maintains superior clean accuracy while achieving outstanding attack success rates, including scenarios with highly concealed single-trigger nodes. Our work highlights critical backdoor risks in web-deployed LM-empowered GFMs and contributes to the development of more robust supervision mechanisms for open-source platforms in the era of foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。