arXiv:2502.02696cs.CL2025-02NAACL被引 4

测试11个大模型对社会规范的认知差异,发现年轻高收入群体更贴近人类共识。

How Inclusively do LMs Perceive Social and Moral Norms?

  • 用规则口诀提示词对比模型与100名人类标注者的判断
  • 发现年轻、高收入群体的模型响应更接近人类共识
  • 适合关注AI伦理与价值观对齐的研究者阅读

语言模型广泛用于决策系统与交互助手,但其道德与社会规范判断是否能反映人类价值观多样性仍不明确。本文研究11个大模型在性别、年龄、收入等群体上的规范认知差异。通过规则口诀(RoTs)提示,将模型输出与100名人类标注者的结果进行对比,并提出绝对距离对齐度量(ADA-Met)评估序数问题的一致性。结果发现模型响应存在显著差异,年轻、高收入群体的判断更贴近人类共识,边缘群体视角代表性不足。研究呼吁加强模型对多元价值观的包容性。代码与提示语已开源,许可协议为CC BY-NC 4.0。

原文摘要 · Abstract (English)

This paper discusses and contains offensive content. Language models (LMs) are used in decision-making systems and as interactive assistants. However, how well do these models making judgements align with the diversity of human values, particularly regarding social and moral norms? In this work, we investigate how inclusively LMs perceive norms across demographic groups (e.g., gender, age, and income). We prompt 11 LMs on rules-of-thumb (RoTs) and compare their outputs with the existing responses of 100 human annotators. We introduce the Absolute Distance Alignment Metric (ADA-Met) to quantify alignment on ordinal questions. We find notable disparities in LM responses, with younger, higher-income groups showing closer alignment, raising concerns about the representation of marginalized perspectives. Our findings highlight the importance of further efforts to make LMs more inclusive of diverse human values. The code and prompts are available on GitHub under the CC BY-NC 4.0 license.

大模型价值观对齐伦理评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。