arXiv:2603.05651cs.CLcs.AI2026-03被引 8

大模型道德判断易受叙述方式和提问形式影响,稳定性堪忧。

The Fragility Of Moral Judgment In Large Language Models

  • 通过扰动故事文本、视角和说服线索,测试模型判断的稳定性。
  • 视角变化导致24.3%判断翻转,远高于表面噪声的7.5%。
  • 评判结果高度依赖提问方式,公平性与可复现性存疑。

人们越来越多地用大语言模型(LLMs)获取日常道德与人际建议,但这些系统无法追问缺失背景,也难以稳定判断道德困境。本文提出一种扰动框架,在保持道德冲突不变的前提下,测试LLM道德判断的稳定性和可操纵性。基于r/AmItheAsshole(2025年1月至3月)的2,939个困境,生成三类扰动:表面编辑(词汇/结构噪声)、视角转换(语气与立场中性化)、说服线索(自我定位、社会证明、模式承认、受害者塑造)。同时对比三种评估协议(输出顺序、指令位置、自由提问)。使用GPT-4.1、Claude 3.7 Sonnet、DeepSeek V3、Qwen2.5-72B四个模型共进行129,156次判断。表面扰动翻转率仅7.5%,接近自一致性噪声范围(4–13%),而视角转换引发24.3%翻转。37.9%的案例在表面噪声下稳定,却在视角变化时翻转,表明模型依赖叙事语气作为实用线索。不明确责任的模糊情境最易波动。说服扰动产生系统性方向偏移。协议选择影响最大:结构化协议间一致率仅67.6%(kappa=0.55),仅35.7%的模型-场景组合在所有协议中一致。结果表明,LLM的道德判断由叙事形式与任务设计共同建构,当结果取决于呈现技巧而非道德实质时,其可复现性与公平性面临严峻挑战。

原文摘要 · Abstract (English)

People increasingly use large language models (LLMs) for everyday moral and interpersonal guidance, yet these systems cannot interrogate missing context and judge dilemmas as presented. We introduce a perturbation framework for testing the stability and manipulability of LLM moral judgments while holding the underlying moral conflict constant. Using 2,939 dilemmas from r/AmItheAsshole (January-March 2025), we generate three families of content perturbations: surface edits (lexical/structural noise), point-of-view shifts (voice and stance neutralization), and persuasion cues (self-positioning, social proof, pattern admissions, victim framing). We also vary the evaluation protocol (output ordering, instruction placement, and unstructured prompting). We evaluated all variants with four models (GPT-4.1, Claude 3.7 Sonnet, DeepSeek V3, Qwen2.5-72B) (N=129,156 judgments). Surface perturbations produce low flip rates (7.5%), largely within the self-consistency noise floor (4-13%), whereas point-of-view shifts induce substantially higher instability (24.3%). A large subset of dilemmas (37.9%) is robust to surface noise yet flips under perspective changes, indicating that models condition on narrative voice as a pragmatic cue. Instability concentrates in morally ambiguous cases; scenarios where no party is assigned blame are most susceptible. Persuasion perturbations yield systematic directional shifts. Protocol choices dominate all other factors: agreement between structured protocols is only 67.6% (kappa=0.55), and only 35.7% of model-scenario units match across all three protocols. These results show that LLM moral judgments are co-produced by narrative form and task scaffolding, raising reproducibility and equity concerns when outcomes depend on presentation skill rather than moral substance.

大模型道德判断可解释性鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。