arXiv:2508.02961cs.AIcs.CL2025-08

让大模型自己识别并防御提示注入攻击,无需外部工具。

Defend LLMs Through Self-Consciousness

  • 利用大模型自身推理能力,通过元认知与仲裁模块实现自我监管。
  • 在多个数据集上测试,部分模型防御成功率接近或达到100%。
  • 轻量高效,适合生成式AI在各类平台落地应用。

本文提出一种新型自意识防御机制,用于对抗大语言模型(LLMs)的提示注入攻击。与依赖外部分类器的传统方法不同,该方法利用大模型自身的推理能力实现自我保护。我们设计了包含元认知与仲裁模块的框架,使模型能够自主评估并调控自身输出。在七个先进大模型上,基于AdvBench和Prompt-Injection-Mixed-Techniques-2024两个数据集进行评估,实验结果表明该方法显著提升了防御成功率,部分模型在增强模式下实现近乎完美(接近100%)的防御效果。同时分析了防御性能提升与计算开销之间的权衡关系。该自意识方法为提升大模型伦理安全性提供了轻量、低成本的解决方案,尤其适用于各类生成式AI应用场景。

原文摘要 · Abstract (English)

This paper introduces a novel self-consciousness defense mechanism for Large Language Models (LLMs) to combat prompt injection attacks. Unlike traditional approaches that rely on external classifiers, our method leverages the LLM's inherent reasoning capabilities to perform self-protection. We propose a framework that incorporates Meta-Cognitive and Arbitration Modules, enabling LLMs to evaluate and regulate their own outputs autonomously. Our approach is evaluated on seven state-of-the-art LLMs using two datasets: AdvBench and Prompt-Injection-Mixed-Techniques-2024. Experiment results demonstrate significant improvements in defense success rates across models and datasets, with some achieving perfect and near-perfect defense in Enhanced Mode. We also analyze the trade-off between defense success rate improvement and computational overhead. This self-consciousness method offers a lightweight, cost-effective solution for enhancing LLM ethics, particularly beneficial for GenAI use cases across various platforms.

大模型安全提示攻击自意识

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。