大模型在道德判断上常自相矛盾,同一问题换种说法答案就变。
Incoherent by Design? On the Moral Self-Consistency of LLMs
- 用三种伦理学框架设计相同情境不同表述的测试题
- 多款大模型在同一流派内判断矛盾率高达78%
- 发现模型无法保持自身道德立场一致,影响对齐效果
大型语言模型在道德敏感场景中的应用日益广泛,但其是否能在不同情境下保持伦理原则的一致性仍不明确。一个能陈述道德原则的模型,在同一情境被重新表述或重构后,可能违背该原则。这种不一致性严重影响依赖模型输出进行道德决策系统的可信度。为探究此问题,我们在德性论、功利主义和义务论三大哲学体系下,构建了若干道德等价情景,仅改变表述方式以反映不同伦理立场与风格扰动。通过评估GPT、Mistral和Llama等多模型输出,我们将回答转化为结构化逻辑语句,识别同一伦理流派内部的矛盾。结果显示,模型在不同情境下的判断矛盾率最高达78%。这揭示了生成式AI中存在普遍的元认知不稳定性,即模型难以维持自身先前输出的一致性。此类不稳定性具有现实后果:当生成系统影响人类信念形成、行为评判与价值吸收时,其不一致会反过来塑造人类推理与决策。若系统无法稳定表达自身的规范承诺,则价值对齐便成为不断变化的目标。因此,我们主张,揭示内部不一致是实现人工智能对齐的必要前提。
原文摘要 · Abstract (English)
LLMs are increasingly used in morally sensitive contexts, yet it is unclear whether they apply ethical principles consistently across situations. A model that can state a moral principle may still violate it when the same scenario is rephrased or reframed. This inconsistency is a problem for any system whose outputs are used to inform moral decisions. If generative systems exhibit internal inconsistency, then the epistemic integrity of AI-mediated systems becomes uncertain. To study this concern, we investigate the stability of moral reasoning in LLMs within a controlled prompting framework across three major philosophical schools of thought: deontology, utilitarianism, and virtue ethics. We construct sets of morally equivalent scenarios in which the underlying situation is held constant while the framing varies to reflect different ethical stances and stylistic perturbations. We then evaluate responses from multiple models, including GPT, Mistral, and Llama. To assess consistency, we convert model outputs into structured logical statements and identify contradictions across responses generated within the same school of thought. Our results reveal substantial inconsistency with contradiction rates reaching up to 78% across scenarios. These findings point to a broader phenomenon of epistemic instability in generative AI wherein models fail to reliably maintain coherence with respect to their own prior outputs. This kind of instability carries real consequences. As generative systems influence how people form beliefs, judge actions, and absorb values, their inconsistencies can shape human reasoning and decision-making as well. Moreover, if a system cannot consistently represent its own normative commitments, then value alignment becomes a moving target rather than a well-defined objective. Thus, we argue that demonstrating internal incoherence is a necessary precursor to AI alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。