大模型知道对错却未必照做,展现人类式认知行为脱节
Knowing But Not Doing: Convergent Morality and Divergent Action in LLMs
- 构建3000条真实场景对话数据集,测试模型在十种价值观下的决策
- 模型间决策高度一致(相关系数≈1.0),但与自身陈述值匹配度低(仅0.3)
- 适合关注AI伦理、价值对齐及认知行为差异的研究者阅读
价值对齐是实现安全且符合社会规范的人工智能的核心。然而,大型语言模型在真实决策情境中如何表征与践行人类价值观仍缺乏研究。本文构建了ValAct-15k数据集,包含3,000个来自Reddit的求助场景,用于激发舒瓦茨基本人类价值观理论定义的十类价值观。通过情景问答与传统问卷两种方式,评估十款前沿大模型(五家美国公司、五家中国公司)及55名人类参与者。结果发现,模型间情景决策几乎完全一致(皮尔逊相关系数≈1.0),而人类个体间差异显著(相关系数区间[-0.79, 0.98])。但无论人类还是模型,自我报告与实际行为间的对应关系均较弱(相关系数分别为0.4和0.3),揭示出系统性的“知行不一”现象。当被指令‘坚持’某一特定价值观时,模型表现下降最多达6.6%,表明其存在角色扮演抗拒。研究显示,尽管对齐训练使模型在价值观上趋于一致,却未能消除类似人类的认知—行为不一致性。
原文摘要 · Abstract (English)
Value alignment is central to the development of safe and socially compatible artificial intelligence. However, how Large Language Models (LLMs) represent and enact human values in real-world decision contexts remains under-explored. We present ValAct-15k, a dataset of 3,000 advice-seeking scenarios derived from Reddit, designed to elicit ten values defined by Schwartz Theory of Basic Human Values. Using both the scenario-based questions and the traditional value questionnaire, we evaluate ten frontier LLMs (five from U.S. companies, five from Chinese ones) and human participants ($n = 55$). We find near-perfect cross-model consistency in scenario-based decisions (Pearson $r \approx 1.0$), contrasting sharply with the broad variability observed among humans ($r \in [-0.79, 0.98]$). Yet, both humans and LLMs show weak correspondence between self-reported and enacted values ($r = 0.4, 0.3$), revealing a systematic knowledge-action gap. When instructed to "hold" a specific value, LLMs' performance declines up to $6.6%$ compared to merely selecting the value, indicating a role-play aversion. These findings suggest that while alignment training yields normative value convergence, it does not eliminate the human-like incoherence between knowing and acting upon values.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。