arXiv:2602.09416cs.CLcs.CY2026-02被引 3

研究发现语言模型会受无关因素影响,道德判断不稳。

Are Language Models Sensitive to Morally Irrelevant Distractors?

  • 用心理学中的无关刺激干扰模型判断
  • 干扰导致判断偏差超30%,即使场景明确
  • 适合关注AI伦理与人类价值观对齐的研究者

随着大语言模型在高风险场景中的广泛应用,确保其行为符合人类价值观变得愈发重要。现有道德评估基准通常通过价值陈述、道德情境或心理问卷来测试模型,隐含假设是模型具有相对稳定的道德偏好。然而,道德心理学研究表明,人类的道德判断也会受到气味、噪音等无关情境因素的影响,挑战了道德稳定性理论。本文借鉴这一“情境主义”视角,构建了一个包含60个无道德相关性的多模态道德干扰项的新数据集,来源于已有情感化图像与叙事的心理学数据集。将这些干扰项注入现有道德基准后,我们发现即便在明确的情境中,干扰项仍能使语言模型的道德判断产生超过30%的偏移,揭示了模型道德判断的不稳定性,凸显了需采用更注重上下文的AI对齐方法。

原文摘要 · Abstract (English)

With the rapid uptake of large language models (LLMs) across high-stakes settings, it is becoming increasingly important to ensure that LLMs behave in ways that align with human values. Existing moral benchmarks for this purpose often prompt LLMs with value statements, moral scenarios, or psychological questionnaires, with the implicit underlying assumption that LLMs report somewhat stable moral preferences. However, moral psychology research has shown that even human moral judgements are sensitive to morally irrelevant situational factors such as the smell of cinnamon rolls or the level of ambient noise, thereby challenging moral theories which assume that human moral judgements are stable. Here we draw inspiration from this "situationist" view of moral psychology to evaluate whether LLMs exhibit similar cognitive moral biases. We curate a novel multimodal dataset of 60 "moral distractors" from existing psychological datasets of emotionally-valenced images and narratives, which have no moral relevance to the situation presented. After injecting these distractors into existing moral benchmarks, we find that moral distractors can shift the moral judgements of LLMs by over 30% even in unambiguous scenarios, highlighting the instability of LLMs' moral judgements and the need for more contextual approaches to AI alignment.

AI伦理道德判断语言模型情境敏感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。