arXiv:2606.04612cs.CL2026-06

融合熵、不确定性与几何特征,提升大模型抗幻觉和对抗攻击能力。

Hybrid Adversarial Defence for Natural Language Understanding Tasks

论文配图:Hybrid Adversarial Defence for Natural Language Understanding Tasks
图 1 · 摘自论文原文
  • 综合熵、不确定性和几何特征构建混合防御框架。
  • 在多个数据集上准确率提升最高达43.34%,攻击成功率降低62.27%。
  • 适用于领域内与域外任务,对提示注入和越狱攻击均有强防御效果。

大语言模型易受幻觉和对抗攻击影响,尽管二者密切相关,现有防御方法通常分别应对。本文提出一种混合防御框架,结合基于熵的模型(抑制幻觉)、基于不确定性的模型和基于几何的模型(降低脆弱性)。在自然语言理解数据集(FEVER、HotpotQA、CSQA、SIQA)上的域内测试中,该模型在干净任务性能上最高提升43.34%准确率,在对抗鲁棒性上最高提升64.92%准确率,攻击成功率降低62.27%。在域外数据集(AeroEngQA、CPIQA)上,对抗鲁棒性同样表现优异(准确率最高提升57.14%)。在提示注入(SafeGuard)和越狱检测(AdvBench、DAN)任务中,攻击成功率相比最先进基线模型最高降低51%。结果表明,联合使用熵、不确定性和几何特征,比单一特征更有效,适用于域内与域外任务。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are vulnerable both to hallucination and adversarial manipulation. Although these problems are closely related, existing defences typically address them separately. We investigate a hybrid defence framework that combines entropy-based models, designed to reduce hallucinations, with uncertainty-based models and geometric-based models, designed to reduce vulnerability. Under in-domain tests on Natural Language Understanding datasets (FEVER, HotpotQA, CSQA, SIQA) we find our hybrid model improves both clean-task performance (up to 43.34\% increase in accuracy) and adversarial robustness (up to 64.92\% improvement in accuracy and 62.27\% reduction in attack success rate). For out-of-distribution datasets (AeroEngQA, CPIQA) we see similar adversarial robustness from our hybrid model (up to 57.14\% improvement in accuracy). For prompt injection (SafeGuard) and jailbreak detection (AdvBench, DAN) datasets our hybrid model is also very strong (up to 51\% reduction in attack success rate compared to state of the art baseline models). Overall, our results show that combining entropy, uncertainty and geometric features provides a more effective defence strategy than using any single feature alone for both in-domain and out-of-distribution tasks.

大模型防御对抗攻击幻觉抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。