arXiv:2601.09724cs.CLcs.AI2026-01被引 2

发现大模型在否定句式下决策易出错,提出可量化评估的鲁棒性框架。

Syntactic Framing Fragility: An Audit of Robustness in LLM Ethical Decisions

  • 通过语法等价变换隔离语义干扰,量化模型在不同表述下的决策一致性。
  • 23个模型测试显示开源模型脆弱性是商用模型的2.2倍,否定句导致80%-97%错误率。
  • 提供符合欧盟AI法案的可落地评估清单,适合高风险场景部署审计。

大型语言模型存在系统性否定敏感问题,但缺乏可规模化部署的评估框架。本文提出语法框架脆弱性(SFF)方法,通过逻辑极性归一化分离语法影响,实现正负表述间的公平对比,并引入语法变异指数(SVI)作为可集成于CI/CD的鲁棒性度量。在14个高风险场景中对23个模型进行审计(共39,975次决策),首次量化了此前仅定性描述的现象,发现开源模型的脆弱性为商用模型的2.2倍。含否定的语法是主要失效模式,部分模型在要求‘代理不行动’时仍以80%-97%概率支持行动。该现象与已有研究中否定抑制失败一致,链式思维推理虽能降低部分模型脆弱性,但效果不一。研究提供了分场景风险画像及符合欧盟AI法案与NIST RMF标准的操作检查表。代码、数据与场景将在发表后公开。

原文摘要 · Abstract (English)

Large language models exhibit systematic negation sensitivity, yet no operational framework exists to measure this vulnerability at deployment scale, especially in high-stakes decisions. We introduce Syntactic Framing Fragility (SFF), a framework for quantifying decision consistency under logically equivalent syntactic transformations. SFF isolates syntactic effects via Logical Polarity Normalization, enabling direct comparison across positive and negative framings while controlling for polarity inversion, and provides the Syntactic Variation Index (SVI) as a robustness metric suitable for CI/CD integration. Auditing 23 models across 14 high-stakes scenarios (39,975 decisions), we establish ground-truth effect sizes for a phenomenon previously characterized only qualitatively and find that open-source models exhibit $2.2x higher fragility than commercial counterparts. Negation-bearing syntax is the dominant failure mode, with some models endorsing actions at 80-97% rates even when asked whether agents not act. These patterns are consistent with negation suppression failure documented in prior work, with chain-of-thought reasoning reducing fragility in some but not all cases. We provide scenario-stratified risk profiles and offer an operational checklist compatible with EU AI Act and NIST RMF requirements. Code, data, and scenarios will be released upon publication.

大模型安全伦理决策鲁棒性评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。