arXiv:2412.06181cs.CRcs.AI2024-12

用递归简化提示,提升大模型抗恶意攻击能力

Enhancing Adversarial Resistance in LLMs with Recursion

  • 通过递归简化复杂提示,增强模型对恶意输入的识别
  • 可有效检测并防御越狱和对抗性提示攻击
  • 适合关注AI安全与模型鲁棒性的研究者使用

大型语言模型(LLMs)在社会中的日益普及,要求其具备应对越狱和对抗性提示漏洞的强健防御能力。本项目提出一种递归框架,利用提示简化技术增强LLM对操纵行为的抵抗能力。通过提高复杂且混淆的对抗性提示的透明度,该方法能更可靠地检测和阻止恶意输入。研究成果旨在解决人工智能安全与防护中的关键问题,为开发能够区分无害输入与含恶意意图提示的系统奠定基础。随着LLMs在多样化应用场景中的持续使用,此类防护机制的重要性将不断上升。

原文摘要 · Abstract (English)

The increasing integration of Large Language Models (LLMs) into society necessitates robust defenses against vulnerabilities from jailbreaking and adversarial prompts. This project proposes a recursive framework for enhancing the resistance of LLMs to manipulation through the use of prompt simplification techniques. By increasing the transparency of complex and confusing adversarial prompts, the proposed method enables more reliable detection and prevention of malicious inputs. Our findings attempt to address a critical problem in AI safety and security, providing a foundation for the development of systems able to distinguish harmless inputs from prompts containing malicious intent. As LLMs continue to be used in diverse applications, the importance of such safeguards will only grow.

大模型安全对抗攻击提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。