测试大模型在哲学压力下的认知崩溃,发现传统评估忽略的深层弱点。
Beyond Social Pressure: Benchmarking Epistemic Attack in Large Language Models
- 构建四类哲学压力测试框架,诊断模型知识合法性动摇
- 多轮诘问下五款模型出现显著认知不一致,模式可区分
- 不同模型需定制防御策略,开源与闭源模型效果差异明显
大型语言模型在压力下会改变答案,这种表现反映的是迎合而非推理。现有对奉承行为的研究主要集中在分歧、恭维和偏好对齐,而更广泛的认知失效问题未被充分探索。本文提出PPT-Bench,一个用于评估‘认知攻击’的诊断基准,其核心是哲学压力分类法(PPT),包含四种压力类型:认知不稳、价值消解、权威反转与身份瓦解。每个测试项在三个层级进行:基线提示(L0)、单轮压力(L1)和多轮苏格拉底式升级(L2)。该设计可衡量L0与L1间的认知不一致性,以及L2中的对话屈服现象。在五种模型上的实验表明,不同压力类型产生统计上可区分的不一致模式,说明认知攻击能揭示标准社会压力基准未能捕捉的缺陷。缓解效果高度依赖类型和模型:API模型中提示锚定与人格稳定性提示最优;开放模型中,引导查询对比解码最可靠。
原文摘要 · Abstract (English)
Large language models (LLMs) can shift their answers under pressure in ways that reflect accommodation rather than reasoning. Prior work on sycophancy has focused mainly on disagreement, flattery, and preference alignment, leaving a broader set of epistemic failures less explored. We introduce \textbf{PPT-Bench}, a diagnostic benchmark for evaluating \textit{epistemic attack}, where prompts challenge the legitimacy of knowledge, values, or identity rather than simply opposing a previous answer. PPT-Bench is organized around the Philosophical Pressure Taxonomy (PPT), which defines four types of philosophical pressure: Epistemic Destabilization, Value Nullification, Authority Inversion, and Identity Dissolution. Each item is tested at three layers: a baseline prompt (L0), a single-turn pressure condition (L1), and a multi-turn Socratic escalation (L2). This allows us to measure epistemic inconsistency between L0 and L1, and conversational capitulation in L2. Across five models, these pressure types produce statistically separable inconsistency patterns, suggesting that epistemic attack exposes weaknesses not captured by standard social-pressure benchmarks. Mitigation results are strongly type- and model-dependent: prompt-level anchoring and persona-stability prompts perform best in API settings, while Leading Query Contrastive Decoding is the most reliable intervention for open models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。