检测语言模型解释中的后门陷阱,无需触发数据也能发现隐藏攻击。
When Explanations Betray Backdoors: Black-Box Auditing for Language Model Classifiers
- 用解释内容的连贯性评估模型是否被植入后门。
- 在五个数据集上实现高于90%的检测准确率,误报率仅5%。
- 适合安全审计人员和模型部署者使用,尤其关注可解释性风险。
带有解释的语言模型分类器被用于内容审核、路由、主题筛选和低资源标注。本文研究在防御者仅有干净校准数据而无触发信息的情况下进行黑盒审计,仅能请求模型输出标签及简短理由或引用证据。提出「有根性漂移」(Groundedness Drift)——一种轻量级评分机制,用于衡量答案总结是否仍基于输入内容。在两个7B规模的模型、五个数据集及四类常见非自适应OpenBackdoor风格攻击下,该方法在所有情况下均优于现有检测器,达到更高AUROC并保持更低残余目标攻击成功率(ASR),且在5%清洁样本假阳性率预算内表现稳定。进一步引入「无支持有根性」(Unsupported Groundedness)作为多探针升级方案,以应对解释伪装场景,虽增强信号,但仍无法弥补自适应攻击下的检测差距。
原文摘要 · Abstract (English)
Language model classifiers with explanations are used for moderation, routing, topic triage, and low-resource annotation. We study black-box auditing when the defender has only clean calibration data without trigger information but can ask the classifier for a label plus a short rationale or quoted evidence. We introduce Groundedness Drift, a lightweight score measuring whether the answer summary remains grounded in the input. Across two 7B backbones, five datasets, and four common non-adaptive OpenBackdoor-style attack families, Groundedness Drift achieves higher AUROC and lower residual target ASR than every compared detector in all cases at a nominal 5\% clean-FPR budget. We then evaluate Unsupported Groundedness, a multi-probe escalation for explanation-camouflage stress cases. Unsupported Groundedness improves signals but does not close the adaptive gap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。