让大模型像人一样警惕虚假政策,自动识别并拒绝欺骗性文本。
Safer Policy Compliance with Dynamic Epistemic Fallback
- 基于人类认知防御机制设计动态安全协议
- 在医疗隐私法等真实政策上实现100%检测率
- 适合高风险场景下保障大模型合规性
人类在日常互动中发展出一种称为'认知警觉'的认知防御机制,以应对欺骗和错误信息的风险。受此启发,本文提出动态认知回退(Dynamic Epistemic Fallback, DEF),一种用于提升大模型推理时对恶意篡改政策文本攻击的防御能力的安全协议。DEF通过单句文本提示,引导大模型识别不一致、拒绝合规,并回退到其参数化知识。基于全球公认的法律政策如HIPAA和GDPR的实证评估显示,DEF显著提升了前沿大模型检测并拒绝篡改政策的能力,其中DeepSeek-R1在某设置下达到100%检测率。该研究推动了基于认知机制的防御方法发展,以增强大模型对利用法律文本进行欺骗性攻击的鲁棒性。
原文摘要 · Abstract (English)
Humans develop a series of cognitive defenses, known as epistemic vigilance, to combat risks of deception and misinformation from everyday interactions. Developing safeguards for LLMs inspired by this mechanism might be particularly helpful for their application in high-stakes tasks such as automating compliance with data privacy laws. In this paper, we introduce Dynamic Epistemic Fallback (DEF), a dynamic safety protocol for improving an LLM's inference-time defenses against deceptive attacks that make use of maliciously perturbed policy texts. Through various levels of one-sentence textual cues, DEF nudges LLMs to flag inconsistencies, refuse compliance, and fallback to their parametric knowledge upon encountering perturbed policy texts. Using globally recognized legal policies such as HIPAA and GDPR, our empirical evaluations report that DEF effectively improves the capability of frontier LLMs to detect and refuse perturbed versions of policies, with DeepSeek-R1 achieving a 100% detection rate in one setting. This work encourages further efforts to develop cognitively inspired defenses to improve LLM robustness against forms of harm and deception that exploit legal artifacts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。