发现大模型在道德困境中会因表述正反而立场反转,影响决策可靠性。
Framing Instability in LLM Ethical Stance: Auditing Negation Sensitivity in Moral Dilemmas
- 用正反双语句对测试16个模型,检测其道德判断是否受表述方式影响
- 小模型在否定表述下认同率最高达100%,比肯定表述高出76个百分点
- 提出新指标NSI衡量立场稳定性,提醒勿用简单二选一评估模型伦理
语言模型越来越多被用于伦理决策咨询,但其立场可能因表述框架变化而改变。我们对16个模型在14个高争议道德困境中进行审计,采用极性配对的提问方式(‘他们应该做某事’与‘他们不该做某事’)。理想情况下,对同一行为的判断不应因正反表述而反转,但我们发现系统性偏离:存在整体立场翻转现象,表明伦理判断易受框架不稳定性影响。参数量1-4B的小型开放模型在肯定表述下仅24%支持某行动,而在否定表述下可高达100%,最高波动达76个百分点。人工标注样本确认该不稳定性真实存在,且二元同意/不同意代理会夸大其程度,说明大模型无法替代人类标注者——因其会隐式合并弃权选项并复制研究中的强制选择偏见。商用模型整体更稳定,但仍有显著波动,跨模型一致性从肯定表述下的73%降至否定表述下的59%。我们主张,因二元选项既夸大了支持度又掩盖了极性依赖性,单语境审计可能误判模型伦理立场,因此提出负向敏感性指数(NSI)作为直接测量立场稳定性的补充工具。若模型立场随表述变化,则不可用于任何高风险决策场景。
原文摘要 · Abstract (English)
Language models are increasingly consulted on ethically consequential questions, yet the stance a model expresses may not survive a change in framing. We audit 16 models across 14 ethically fraught dilemmas using polarity-paired proposals ("They should X" / "They should not X"). A model's judgment of the underlying action should not reverse merely because the question is phrased as a prohibition rather than a prescription and yet, we find systematic deviations from this invariance including wholesale endorsement flips, indicating that ethical decisions are vulnerable to framing instability. Small open-weight models (1-4B parameters) endorse a proposed action 24% of the time under affirmative framing but up to 100% under negated framings, a swing of as much as 76 percentage points. Human coding of a response sample confirms the instability is genuine while showing that binary agree/disagree proxies over-state its magnitude, suggesting that an LLM judge cannot replace human coders because it silently collapses abstentions and mirrors the very forced-choice bias under study. Commercial models are for the most part more stable but still shift substantially, with cross-model agreement dropping from 73% on the bare affirmative framing to 59% under simple negation. We argue that because binary agree/disagree formats both inflate apparent endorsement and mask polarity-dependence, single-phrasing audits can misreport a model's ethical stance, and we propose the Negation Sensitivity Index (NSI) as a complement that measures stance stability directly. A model whose stance flips with phrasing cannot be relied upon in any high-stakes decision scenario.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。