arXiv:2412.00956cs.AIcs.CL2024-12被引 6

大模型难以准确反映全球多元道德观,存在文化偏差。

Large Language Models as Mirrors of Societal Moral Standards

  • 用全球40国调查数据对比模型与真实道德观
  • 多语言模型普遍表现不佳,仅BLOOM部分相关
  • 揭示当前AI缺乏跨文化价值观理解能力

先前研究已表明语言模型能在一定程度上反映不同文化背景下的道德规范。本研究旨在复现并进一步检验这些发现,重点关注同性恋和离婚等议题。通过使用涵盖40多个国家的全球价值调查(WVS)与皮尤研究中心(PEW)数据,评估模型在捕捉多元文化道德观方面的有效性。结果表明,单语和多语模型均存在偏见,通常无法准确呈现不同文化的道德复杂性。尽管BLOOM模型表现最佳,显示出一定正相关性,但仍未能实现全面的道德理解。研究强调当前预训练语言模型在处理跨文化价值差异方面的局限性,并呼吁开发更具备文化敏感性的智能系统,以更好契合普世人类价值观。

原文摘要 · Abstract (English)

Prior research has demonstrated that language models can, to a limited extent, represent moral norms in a variety of cultural contexts. This research aims to replicate these findings and further explore their validity, concentrating on issues like 'homosexuality' and 'divorce'. This study evaluates the effectiveness of these models using information from two surveys, the WVS and the PEW, that encompass moral perspectives from over 40 countries. The results show that biases exist in both monolingual and multilingual models, and they typically fall short of accurately capturing the moral intricacies of diverse cultures. However, the BLOOM model shows the best performance, exhibiting some positive correlations, but still does not achieve a comprehensive moral understanding. This research underscores the limitations of current PLMs in processing cross-cultural differences in values and highlights the importance of developing culturally aware AI systems that better align with universal human values.

大模型道德对齐跨文化价值对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。