arXiv:2411.11937cs.LGcs.AI2024-11NeurIPS被引 14

分析三大强化学习人类反馈数据集中的价值观分布,发现知识类价值最突出。

Value Imprint: A Technique for Auditing the Human Values Embedded in RLHF Datasets

  • 构建哲学与伦理学融合的价值分类体系,标注6501条偏好数据
  • 发现知识类价值在三数据集中占比最高,而公正、权利等价值最少
  • 开源数据集供后续研究,助力价值观对齐的可解释性审计

大型语言模型日益通过强化学习人类反馈(RLHF)数据集进行微调以对齐人类偏好与价值观。然而,关于这些数据集中具体嵌入了哪些人类价值观的研究仍十分有限。本文提出Value Imprint框架,用于审计和分类RLHF数据集中的隐含价值观。我们通过两阶段分析,对Anthropic/hh-rlhf、OpenAI WebGPT Comparisons及Alpaca GPT-4-LLM三个数据集进行了实证研究。第一阶段基于哲学、价值论与伦理学文献构建人类价值分类体系,并对6,501条偏好样本进行标注;第二阶段利用标注结果训练基于Transformer的机器学习模型,实现对三数据集的价值审计。结果显示,信息-效用类价值(如智慧/知识、信息获取)在所有数据集中占主导地位,而亲社会与民主类价值(如福祉、正义、人/动物权利)则显著不足。该发现对构建符合社会规范的语言模型具有重要启示。我们已公开相关数据集,支持后续研究。

原文摘要 · Abstract (English)

LLMs are increasingly fine-tuned using RLHF datasets to align them with human preferences and values. However, very limited research has investigated which specific human values are operationalized through these datasets. In this paper, we introduce Value Imprint, a framework for auditing and classifying the human values embedded within RLHF datasets. To investigate the viability of this framework, we conducted three case study experiments by auditing the Anthropic/hh-rlhf, OpenAI WebGPT Comparisons, and Alpaca GPT-4-LLM datasets to examine the human values embedded within them. Our analysis involved a two-phase process. During the first phase, we developed a taxonomy of human values through an integrated review of prior works from philosophy, axiology, and ethics. Then, we applied this taxonomy to annotate 6,501 RLHF preferences. During the second phase, we employed the labels generated from the annotation as ground truth data for training a transformer-based machine learning model to audit and classify the three RLHF datasets. Through this approach, we discovered that information-utility values, including Wisdom/Knowledge and Information Seeking, were the most dominant human values within all three RLHF datasets. In contrast, prosocial and democratic values, including Well-being, Justice, and Human/Animal Rights, were the least represented human values. These findings have significant implications for developing language models that align with societal values and norms. We contribute our datasets to support further research in this area.

价值观审计RLHF数据对齐可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。