arXiv:2412.00962cs.AIcs.CL2024-12被引 6

测试大模型能否反映不同文化间的道德差异与共识。

LLMs as mirrors of societal moral standards: reflection of cultural divergence and agreement across ethical topics

  • 用三种方法比对模型输出与真实调查数据的道德评分差异。
  • 模型在反映跨文化道德差异方面表现普遍较差。
  • 适合关注AI伦理与跨文化公平性的研究者阅读。

大型语言模型(LLMs)因性能提升在多个领域日益重要,但其训练数据带来的性别、种族和文化偏见仍引发关注。本研究探讨这些模型是否能准确反映跨文化道德观念的差异与共识。通过三种方法评估:(1)比较模型生成与调查所得的道德评分方差;(2)聚类一致性分析,检验模型生成的国家聚类与调查数据聚类的对应关系;(3)使用直接对比提示探测模型。所有方法均采用系统化提示和词对设计,以衡量模型对文化道德态度的理解能力。结果表明,所测试模型在反映跨文化道德差异与共识方面整体表现不佳,凸显提升模型捕捉文化细微差别能力的必要性。研究为全球语境下大模型的伦理开发与部署提供参考,强调减少偏见、促进多元文化公平表达的重要性。

原文摘要 · Abstract (English)

Large language models (LLMs) have become increasingly pivotal in various domains due the recent advancements in their performance capabilities. However, concerns persist regarding biases in LLMs, including gender, racial, and cultural biases derived from their training data. These biases raise critical questions about the ethical deployment and societal impact of LLMs. Acknowledging these concerns, this study investigates whether LLMs accurately reflect cross-cultural variations and similarities in moral perspectives. In assessing whether the chosen LLMs capture patterns of divergence and agreement on moral topics across cultures, three main methods are employed: (1) comparison of model-generated and survey-based moral score variances, (2) cluster alignment analysis to evaluate the correspondence between country clusters derived from model-generated moral scores and those derived from survey data, and (3) probing LLMs with direct comparative prompts. All three methods involve the use of systematic prompts and token pairs designed to assess how well LLMs understand and reflect cultural variations in moral attitudes. The findings of this study indicate overall variable and low performance in reflecting cross-cultural differences and similarities in moral values across the models tested, highlighting the necessity for improving models' accuracy in capturing these nuances effectively. The insights gained from this study aim to inform discussions on the ethical development and deployment of LLMs in global contexts, emphasizing the importance of mitigating biases and promoting fair representation across diverse cultural perspectives.

大模型伦理跨文化道德判断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。