语言模型易被'必须''应该'等词误导,误判非义务场景为义务。
Deontological Keyword Bias: The Impact of Modal Expressions on Normative Judgments of Language Models
- 用带'必须''应该'的提示,让模型把90%常识场景判断为义务
- 不同模型、问题类型和回答格式下均出现此偏差
- 通过少量示例+推理提示可有效缓解该偏见
大型语言模型(LLMs)在道德与伦理推理中扮演日益重要的角色,但判断标准对人类而言尚不明确。尽管已有大量关于模型对齐的研究,但对模型如何判断义务的问题仍关注不足。本文揭示了语言模型存在显著的道义关键词偏差(Deontological Keyword Bias, DKB):当提示中包含'必须'或'应当'等情态表达时,模型倾向于将非义务性情境判定为义务性。我们在多种常识场景中发现,加入情态表达后,模型判断为义务的比例超过90%,且该现象在不同模型家族、问题类型和答案格式中均具一致性。为此,我们提出一种结合少量示例与推理提示的判断策略以缓解该偏差。研究揭示了情态表达作为语言框架如何影响模型的规范性决策,强调了纠正此类偏见对实现判断对齐的重要性。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly engaging in moral and ethical reasoning, where criteria for judgment are often unclear, even for humans. While LLM alignment studies cover many areas, one important yet underexplored area is how LLMs make judgments about obligations. This work reveals a strong tendency in LLMs to judge non-obligatory contexts as obligations when prompts are augmented with modal expressions such as must or ought to. We introduce this phenomenon as Deontological Keyword Bias (DKB). We find that LLMs judge over 90\% of commonsense scenarios as obligations when modal expressions are present. This tendency is consist across various LLM families, question types, and answer formats. To mitigate DKB, we propose a judgment strategy that integrates few-shot examples with reasoning prompts. This study sheds light on how modal expressions, as a form of linguistic framing, influence the normative decisions of LLMs and underscores the importance of addressing such biases to ensure judgment alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。