arXiv:2505.19674cs.CL2025-05ACL被引 6

用词关联分析大模型与西方人的道德差异,发现系统性不同。

Comparing Moral Values in Western English-speaking societies and LLMs with Word Associations

  • 通过词关联构建人类与大模型的语义网络
  • 基于道德基础理论传播价值观,揭示深层差异
  • 适合研究模型伦理与跨文化比较的研究者

随着大语言模型影响力扩大,理解其反映的道德价值观变得愈发重要。直接通过提示评估模型道德认知存在挑战:训练数据可能泄露人类规范,且结果对提示设计敏感。为此,我们提出使用词关联——一种已被证明能反映人类道德推理的低层表征——来获得更稳健的大模型道德推理画像。研究对比了以英语为主的西方社会群体与主要在英语数据上训练的大模型之间的道德关联差异。首先,构建了一个大规模的由大模型生成的词关联数据集,类比已有真实人类词关联数据。其次,提出一种新方法,通过从道德基础理论中提取种子词,在人类与大模型生成的关联图谱中传播道德价值。最后,比较所得道德概念化结果,揭示出英语使用者与大模型词关联中存在细致而系统的道德差异。

原文摘要 · Abstract (English)

As the impact of large language models increases, understanding the moral values they reflect becomes ever more important. Assessing the nature of moral values as understood by these models via direct prompting is challenging due to potential leakage of human norms into model training data, and their sensitivity to prompt formulation. Instead, we propose to use word associations, which have been shown to reflect moral reasoning in humans, as low-level underlying representations to obtain a more robust picture of LLMs' moral reasoning. We study moral differences in associations from western English-speaking communities and LLMs trained predominantly on English data. First, we create a large dataset of LLM-generated word associations, resembling an existing data set of human word associations. Next, we propose a novel method to propagate moral values based on seed words derived from Moral Foundation Theory through the human and LLM-generated association graphs. Finally, we compare the resulting moral conceptualizations, highlighting detailed but systematic differences between moral values emerging from English speakers and LLM associations.

道德推理词关联大模型伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。