无需人工标注,用算法从语言模型中挖掘隐藏的道德判断能力
Unsupervised Elicitation of Moral Values from Language Models
- 用无监督算法ICM挖掘预训练模型的内在道德推理能力
- 在多个基准上表现优于人类标注和聊天模型,尤其在公正与常识道德上提升显著
- 可有效降低种族、阶级、政治等社会偏见,适合对齐AI行为的研究者
随着AI系统日益普及,将其行为锚定于人类价值观至关重要。以往研究认为语言模型(LMs)具备有限的内在道德推理能力,需通过显式教学来增强。然而,构建道德评估的黄金标准数据困难,因价值框架多元且普遍存在偏见。本文探索无监督挖掘作为替代方案,检验预训练(基础)模型是否具备可被激发的内在道德推理能力,而无需人类监督。在三个基准数据集和四种语言模型上,采用内部一致性最大化(ICM)算法测试其能否可靠标注道德判断、跨道德框架泛化并缓解社会偏见。结果表明,ICM在Norm Bank和ETHICS基准上均优于所有预训练及聊天模型基线;基于ICM标签微调的表现与或超越人类标注。在理论驱动的道德框架中,ICM在正义和常识道德上获得最大相对提升。尽管聊天模型的社会偏见错误率与预训练模型相当,但ICM使其降低超过一半,尤其在种族、社会经济地位和政治领域改善最明显。研究提示:预训练模型具备可被无监督方法如ICM激发的潜在道德推理能力,为人工智能对齐提供了可扩展路径。
原文摘要 · Abstract (English)
As AI systems become pervasive, grounding their behavior in human values is critical. Prior work suggests that language models (LMs) exhibit limited inherent moral reasoning, leading to calls for explicit moral teaching. However, constructing ground truth data for moral evaluation is difficult given plural frameworks and pervasive biases. We investigate unsupervised elicitation as an alternative, asking whether pretrained (base) LMs possess intrinsic moral reasoning capability that can be surfaced without human supervision. Using the Internal Coherence Maximization (ICM) algorithm across three benchmark datasets and four LMs, we test whether ICM can reliably label moral judgments, generalize across moral frameworks, and mitigate social bias. Results show that ICM outperforms all pre-trained and chatbot baselines on the Norm Bank and ETHICS benchmarks, while fine-tuning on ICM labels performs on par with or surpasses those of human labels. Across theoretically motivated moral frameworks, ICM yields its largest relative gains on Justice and Commonsense morality. Furthermore, although chatbot LMs exhibit social bias failure rates comparable to their pretrained ones, ICM reduces such errors by more than half, with the largest improvements in race, socioeconomic status, and politics. These findings suggest that pretrained LMs possess latent moral reasoning capacities that can be elicited through unsupervised methods like ICM, providing a scalable path for AI alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。