arXiv:2607.16903cs.CYcs.AI2026-07

让生成式AI学会人类价值观的系统方法

A Method for Learning Value Systems in Generative AI

  • 基于成对提示-响应偏好数据,同时学习价值根基与加权价值系统
  • 在多个数据集上表现优于基线,且解释性更强
  • 适合需要可解释价值对齐的生成式AI研究者

具备价值意识的AI系统需要显式的计算化人类价值观表示(即价值根基)及其整合为价值体系,以使决策与人类一致。由于这些表示难以直接获取,价值学习通过观察人类行为来推断。本文针对生成式AI中缺乏具身价值学习方法的问题:现有方法通常仅复制人类偏好,而未考虑价值对齐的多维结构,或缺乏严谨的价值体系推导机制。为此,我们将在先前验证有效的价值体系学习方法适配至生成式AI场景,基于成对提示-响应偏好数据,同时学习:i) 由多目标奖励模型给出的一组价值的实现方式(价值根基),ii) 该根基模型的加权线性标量化形式的价值体系表示。为确保所学价值体系基于连贯的价值表示,算法动态优先推进价值根基学习过程。我们在提示-响应偏好数据集上评估该方法,结果表明其性能具有竞争力,相较于基线方法保持最小权衡,同时显著提升可解释性。

原文摘要 · Abstract (English)

Value-aware AI systems require explicit computational representations of human values (groundings) and their aggregation into value systems in order to align their decisions with ours. As such representations are difficult to elicit, value learning seeks to infer them by observing human behaviour. This work addresses the lack of grounded value learning methods in generative AI: existing approaches typically replicate human preferences without awareness of the multidimensional structure of value alignment, or lack principled value system elicitation methods. To address these gaps, we adapt a previously validated value system learning method to the generative AI setting, which, based on pairwise prompt-response preference data, simultaneously learns: i) an implementation of a grounding for a set of values given by a multi-objective reward model, and ii) a value system representation in the form of a weighted linear scalarization of the previous grounding model. To ensure that the learned value systems are based on coherent value representations, our algorithm dynamically prioritizes the grounding learning process. We evaluate the method against baselines and a contemporary method on prompt-response preference datasets. Results show competitive performance and minimal trade-offs against the baselines, while improving explainability.

价值对齐生成式AI奖励建模可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。