arXiv:2504.20769cs.CLcs.AI2025-04被引 5

用结构化防御推理提升大模型抗干扰能力

Chain-of-Defensive-Thought: Structured Reasoning Elicits Robustness in Large Language Models against Reference Corruption

  • 提供少量带防御性推理的示范样本
  • GPT-4o在参考信息被注入攻击时准确率从3%升至50%
  • 方法简单通用,适合增强模型鲁棒性

思维链提示已证明能有效提升大语言模型的推理能力。本文探索如何利用这种增强的推理能力,来改善大模型在非推理任务中的鲁棒性。我们提出一种名为「链式防御思维」(Chain-of-Defensive-Thought)的简单方法:仅需提供少量带有结构化、防御性推理的示范样本作为提示。实验证明,该方法效果显著,尤其在复杂对抗场景下表现突出。例如,在Natural Questions任务中,当标准提示下10个参考文档中有1个被提示注入攻击破坏时,GPT-4o的准确率从60%暴跌至3%;而使用链式防御思维提示后,准确率仍维持在50%。该方法无需修改模型,且适用范围广。

原文摘要 · Abstract (English)

Chain-of-thought prompting has demonstrated great success in facilitating the reasoning abilities of large language models. In this work, we explore how these enhanced reasoning abilities can be exploited to improve the robustness of large language models in tasks that are not necessarily reasoning-focused. In particular, we show how a wide range of large language models exhibit significantly improved robustness against reference corruption using a simple method called chain-of-defensive-thought, where only a few exemplars with structured and defensive reasoning are provided as demonstrations. Empirically, the improvements can be astounding, especially given the simplicity and applicability of the method. For example, in the Natural Questions task, the accuracy of GPT-4o degrades from 60% to as low as 3% with standard prompting when 1 out of 10 references provided is corrupted with prompt injection attacks. In contrast, GPT-4o using chain-of-defensive-thought prompting maintains an accuracy of 50%.

大模型鲁棒性防御推理提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。