arXiv:2409.13521cs.CLcs.AI2024-09中稿 · publication with A…综述被引 10

梳理大模型与道德基础理论的结合进展,探索让AI更懂人类价值观的方法。

A Survey on Moral Foundation Theory and Pre-Trained Language Models: Current Advances and Challenges

  • 用预训练语言模型分析文本中的道德维度,结合道德基础理论框架。
  • 总结了相关数据集和词典,揭示大模型在道德倾向上的表现差异。
  • 适合关注伦理对齐、可解释AI的研究者或从业者阅读。

道德价值根植于早期文明,以规范和法律形式维系社会秩序与公共利益,对理解人类行为心理和文化取向至关重要。道德基础理论(MFT)是一个成熟框架,识别出不同文化塑造个体与社会生活的核心道德基础。近年来,自然语言处理领域的进展,特别是预训练语言模型(PLMs),使得从文本中提取和分析道德维度成为可能。本综述系统回顾了基于MFT的PLMs研究,分析了PLMs中的道德倾向及其在MFT背景下的应用。同时,我们梳理了相关数据集和词典,讨论了当前趋势、局限与未来方向。通过构建PLMs与MFT交叉领域的结构化图景,本文连接了道德心理学洞见与大模型研究,为开发具备道德意识的AI系统奠定基础。

原文摘要 · Abstract (English)

Moral values have deep roots in early civilizations, codified within norms and laws that regulated societal order and the common good. They play a crucial role in understanding the psychological basis of human behavior and cultural orientation. The Moral Foundation Theory (MFT) is a well-established framework that identifies the core moral foundations underlying the manner in which different cultures shape individual and social lives. Recent advancements in natural language processing, particularly Pre-trained Language Models (PLMs), have enabled the extraction and analysis of moral dimensions from textual data. This survey presents a comprehensive review of MFT-informed PLMs, providing an analysis of moral tendencies in PLMs and their application in the context of the MFT. We also review relevant datasets and lexicons and discuss trends, limitations, and future directions. By providing a structured overview of the intersection between PLMs and MFT, this work bridges moral psychology insights within the realm of PLMs, paving the way for further research and development in creating morally aware AI systems.

道德对齐大模型心理学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。