测试大模型是理解深层价值还是仅模仿表面偏好。
Deep Value Benchmark: Measuring Whether Models Generalize Deep Values or Shallow Preferences
- 设计对照实验,分离深层价值与表层特征进行测试。
- 9个模型平均深层价值泛化率仅0.30,低于随机水平。
- 适合关注AI对齐、价值观泛化能力的研究者阅读。
我们提出深度价值基准(DVB),一种评估大语言模型是否真正学习人类深层价值观而非仅模仿表面偏好。该区分对人工智能对齐至关重要:掌握深层价值的系统更可能稳健地泛化人类意图,而仅依赖偏好数据中表层模式的系统则易产生不一致行为。DVB采用新型实验设计,在训练阶段人为制造深层价值(如道德原则)与浅层特征(如语言形式)之间的共现关联——例如用户始终偏好(非伤害性,正式语言)而非(公正性,非正式语言)选项。测试阶段打破这种关联,呈现(公正性,正式语言)与(非伤害性,非正式语言)的选择。该设计可精确测量模型的深层价值泛化率(DVGR),即基于深层价值而非浅层特征进行推理的概率。在9个不同模型中,平均DVGR仅为0.30,所有模型的深层价值泛化率均低于随机水平;更大模型的DVGR反而略低。我们已公开数据集,并经过三次独立的人类验证。DVB为对齐的核心特性提供了可解释的度量方式。
原文摘要 · Abstract (English)
We introduce the Deep Value Benchmark (DVB), an evaluation framework that directly tests whether large language models (LLMs) learn fundamental human values or merely surface-level preferences. This distinction is critical for AI alignment: Systems that capture deeper values are likely to generalize human intentions robustly, while those that capture only superficial patterns in preference data risk producing misaligned behavior. The DVB uses a novel experimental design with controlled confounding between deep values (e.g., moral principles) and shallow features (e.g., superficial attributes). In the training phase, we expose LLMs to human preference data with deliberately correlated deep and shallow features -- for instance, where a user consistently prefers (non-maleficence, formal language) options over (justice, informal language) alternatives. The testing phase then breaks these correlations, presenting choices between (justice, formal language) and (non-maleficence, informal language) options. This design allows us to precisely measure a model's Deep Value Generalization Rate (DVGR) -- the probability of generalizing based on the underlying value rather than the shallow feature. Across 9 different models, the average DVGR is just 0.30. All models generalize deep values less than chance. Larger models have a (slightly) lower DVGR than smaller models. We are releasing our dataset, which was subject to three separate human validation experiments. DVB provides an interpretable measure of a core feature of alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。