提出化学合规的后门攻击检测与构造方法,揭示真实分子数据管道中的安全漏洞。
Rethinking Molecular Graph Backdoors under Chemistry-aware Admission

- 引入化学合规检查协议ChemGuard,验证分子数据是否可进入真实学习流程
- 设计无需模型信息的ChemBack攻击,实现高成功率且全合规的后门注入
- 适用于分子机器学习安全评估,尤其关注药物研发与化学数据治理场景
分子图神经网络的后门攻击通常在抽象图编辑层面评估,但真实分子学习流程需经过解析、清洗、标准化及图-字符串一致性校验。本文提出ChemGuard协议,用于测试分子记录能否通过现实数据管道的准入检验——仅当其分子字符串可清洗且重建图与提交图一致时才被接纳。在此视角下,许多现有图级后门因化学无效或表示不一致而失效。进一步发现,仅靠准入检查不足以防御后门。为此提出ChemBack:一种基于指纹相似度排序的无模型后门攻击,利用化学结构、标签、指纹和公开有效性检查构建合法毒药,无需目标模型、代理模型、梯度或训练代码。在多个分子基准、验证器、架构与防御下,ChemBack均实现高攻击成功率且毒药完全合规,同时保持干净准确率。结果表明,化学合规机制虽能抑制多数图级后门,但化学有效且目标对齐的后门仍构成现实威胁。
原文摘要 · Abstract (English)
Backdoor attacks on molecular graph neural networks (GNNs) are typically evaluated as abstract graph edits, but real molecular learning pipelines do not train on arbitrary graphs. Molecular records must first survive parsing, sanitization, canonicalization, and graph-string consistency checks. We formalize this overlooked admission stage as ChemGuard, an operational protocol for testing whether a submitted molecular record can enter a realistic learning pipeline, while complementing existing defenses. ChemGuard admits a record only when its molecular string is sanitizable and the graph reconstructed from that string matches the submitted molecular graph. Under this operational view, many existing graph-based backdoors lose much of their apparent efficacy because their poisons are chemically invalid or representation-inconsistent. We then show that admission checks alone are insufficient to rule out molecular backdoors. We propose ChemBack, an admission-aware molecular backdoor attack that constructs chemically feasible motif-anchor attachments and ranks admitted candidates by fingerprint-based Tanimoto similarity to clean target-class molecules. ChemBack is model-free during trigger selection, using molecular structures, target labels, fingerprints, and public validity checks, but no victim model, surrogate GNN, learned embedding, gradient, logit, or training-code access. Across molecular benchmarks, validators, architectures, and defenses, \textbf{ChemBack} achieves high attack success with fully admitted poisons while preserving clean accuracy. Our results reveal a two-sided lesson, chemistry-aware admission suppresses many graph-only backdoors, yet chemically valid and target-aligned molecular backdoors remain a practical threat.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。