arXiv:2602.03160cs.AIcs.CL2026-02中稿 · ICML

首个可调控强度的多价值对齐框架,让大模型更精准响应人类价值观。

VALUEFLOW: Toward Pluralistic and Steerable Value-based Alignment in Large Language Models

  • 构建分层价值空间与强度数据库,统一提取、评估与控制
  • 在10个模型、4种理论下发现价值控制存在不对称性
  • 适合研究价值对齐、可控生成与伦理可控性的学者

将大型语言模型(LLMs)与人类多元价值观对齐仍是核心挑战:基于偏好的方法常无法捕捉深层动机原则。基于价值的方法提供更严谨路径,但仍有三大缺陷:提取忽略层级结构,评估仅检测存在而无强度校准,且模型在可控强度下的可调性仍不充分。为此,我们提出VALUEFLOW——首个涵盖提取、评估与调控且具备强度校准能力的统一框架。该框架包含三个组件:(i) HIVES,一种分层价值嵌入空间,用于捕捉跨理论与内部的价值结构;(ii) Value Intensity DataBase (VIDB),一个大规模标注文本库,其强度估计通过排序聚合获得;(iii) 基于锚点的评估器,通过将模型输出与VIDB样本对比排序,生成一致的强度评分。利用VALUEFLOW,我们在十个模型和四种价值理论下开展大规模研究,揭示了可调性中的不对称性,并发现多价值控制的组合规律。本工作建立了一套可扩展的价值强度评估与控制基础设施,推动了大模型的多元价值观对齐。

原文摘要 · Abstract (English)

Aligning Large Language Models (LLMs) with the diverse spectrum of human values remains a central challenge: preference-based methods often fail to capture deeper motivational principles. Value-based approaches offer a more principled path, yet three gaps persist: extraction often ignores hierarchical structure, evaluation detects presence but not calibrated intensity, and the steerability of LLMs at controlled intensities remains insufficiently understood. To address these limitations, we introduce VALUEFLOW, the first unified framework that spans extraction, evaluation, and steering with calibrated intensity control. The framework integrates three components: (i) HIVES, a hierarchical value embedding space that captures intra- and cross-theory value structure; (ii) the Value Intensity DataBase (VIDB), a large-scale resource of value-labeled texts with intensity estimates derived from ranking-based aggregation; and (iii) an anchor-based evaluator that produces consistent intensity scores for model outputs by ranking them against VIDB panels. Using VALUEFLOW, we conduct a comprehensive large-scale study across ten models and four value theories, identifying asymmetries in steerability and composition laws for multi-value control. This paper establishes a scalable infrastructure for evaluating and controlling value intensity, advancing pluralistic alignment of LLMs.

价值对齐可控生成大模型伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。