用智能代理模拟攻击者,让大模型在保持意图下产生事实错误。
Dive into Ambiguity: A*-Inspired Multi-Agents Commonsense Obfuscation Attack on LLM Prompts

- 基于A*启发式搜索,分阶段动态调整改写强度,先保守后激进。
- 在多个大模型上成功率超传统方法,且所需尝试次数更少。
- 可解释攻击机制,适合研究模型安全与对抗攻防的学者。
大型语言模型(LLMs)在推理和知识密集型任务中表现优异,但对提示层面的对抗攻击仍脆弱——此类攻击在保留原始意图的同时引发常识性幻觉。这一漏洞尤为紧迫,因大模型正快速应用于安全关键领域,事实可靠性不容妥协。现有攻击方法或效率低下,或无法捕捉真实攻击者的自适应策略。本文提出一种受A*启发的事实错误诱导框架,用于生成语义一致但被混淆的提示。核心是基于动态语义发散系数 $γ$ 的分层重写策略,遵循逆向模拟退火过程:初期保守修改,后期激进混淆。为增强可解释性,进一步引入代理机制标注,发现并优化对抗机制,实现可解释的逆向优化。理论上,证明了提示重写遵循压缩递推关系,当 $γ$ 减小时导致语义坍缩。实验表明,在多种大模型上,该方法在更高攻击成功率下,仅需更少尝试次数,兼具高效性与有效性。
原文摘要 · Abstract (English)
Large language models (LLMs) excel in reasoning and knowledge-intensive tasks but remain vulnerable to prompt-level adversarial attacks that preserve intent while triggering commonsense hallucinations. This vulnerability is urgent, as LLMs are rapidly integrated into safety-critical domains where factual reliability is non-negotiable. Existing attack methods either lack efficiency or fail to capture the adaptive strategies of real-world adversaries. We propose an A*-inspired Factual Error Induction Framework, a framework for generating semantically aligned yet obfuscated prompts. At its core is a Hierarchical Rewrite Strategy guided by a dynamic semantic dispersion coefficient $γ$ that balances conservative edits early with aggressive obfuscations later, following a reverse simulated annealing schedule. To enhance interpretability, we further introduce Agentic Mechanism Labeling, which discovers and refines adversarial mechanisms, offering interpretable reverse optimization. Theoretically, we prove that prompt rewriting follows a contractive recurrence, leading to semantic collapse as $γ$ decreases. Empirically, across diverse LLMs, our method achieves higher attack success rates than exhaustive exploration while requiring fewer attempts, demonstrating both efficiency and effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。