大模型更关注道德判断,忽视常识矛盾。
Common Sense vs. Morality: The Curious Case of Narrative Focus Bias in LLMs
- 设计新数据集CoMoral,测试模型在道德困境中的常识识别能力。
- 模型在无提示时几乎无法发现常识矛盾,且对次要角色更敏感。
- 适合研究模型推理偏差、伦理对齐的学者参考。
大型语言模型(LLMs)在各类实际应用中日益普及,需兼具道德合理性与知识感知能力。本文揭示当前模型的一个关键缺陷:过度侧重道德推理而忽略常识理解。为此,我们构建了CoMoral——一个包含常识矛盾的道德困境新基准数据集。通过对十种不同规模的LLMs进行评估,发现现有模型在缺乏显式提示时普遍难以识别此类矛盾。进一步观察到一种普遍存在的叙事焦点偏差:当常识矛盾出现在次要角色而非主要叙述者身上时,模型更易察觉。本研究强调,需通过增强推理感知训练来提升大模型的常识鲁棒性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly deployed across diverse real-world applications and user communities. As such, it is crucial that these models remain both morally grounded and knowledge-aware. In this work, we uncover a critical limitation of current LLMs -- their tendency to prioritize moral reasoning over commonsense understanding. To investigate this phenomenon, we introduce CoMoral, a novel benchmark dataset containing commonsense contradictions embedded within moral dilemmas. Through extensive evaluation of ten LLMs across different model sizes, we find that existing models consistently struggle to identify such contradictions without prior signal. Furthermore, we observe a pervasive narrative focus bias, wherein LLMs more readily detect commonsense contradictions when they are attributed to a secondary character rather than the primary (narrator) character. Our comprehensive analysis underscores the need for enhanced reasoning-aware training to improve the commonsense robustness of large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。