构建评估大模型价值观的综合性平台,突破安全偏见局限。
Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values
- 以基本价值为基底,从动机维度全面刻画模型内在价值观
- 采用动态生成测试题,实现对模型行为与价值一致性的实时评估
- 融合多元文化价值观权重,量化模型在不同价值维度的对齐程度
随着大语言模型取得显著进展,使其价值观与人类保持一致已成为负责任发展和定制化应用的迫切需求。然而,现有评估方法难以同时满足三大目标:(1)价值清晰性——当前评估多聚焦于偏见、毒性等安全风险,未能全面揭示模型深层价值观;(2)评估有效性——现有静态开源基准易受数据污染,且仅检验模型对价值观的知识掌握,而非真实行为一致性;(3)价值多样性——忽视个体与文化间的价值差异。为此,本文提出价值罗盘评测平台,包含三个模块:(i)基于动机区分的基本价值体系,实现对模型价值观的全局刻画;(ii)采用生成式演化评估框架,通过自适应测试项动态追踪模型演化,并直接评估其在真实场景中的行为契合度;(iii)设计加权多维指标,将特定价值对齐程度量化为多个维度的加权和,权重由多元价值共识确定。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) achieve remarkable breakthroughs, aligning their values with humans has become imperative for their responsible development and customized applications. However, there still lack evaluations of LLMs values that fulfill three desirable goals. (1) Value Clarification: We expect to clarify the underlying values of LLMs precisely and comprehensively, while current evaluations focus narrowly on safety risks such as bias and toxicity. (2) Evaluation Validity: Existing static, open-source benchmarks are prone to data contamination and quickly become obsolete as LLMs evolve. Additionally, these discriminative evaluations uncover LLMs' knowledge about values, rather than valid assessments of LLMs' behavioral conformity to values. (3) Value Pluralism: The pluralistic nature of human values across individuals and cultures is largely ignored in measuring LLMs value alignment. To address these challenges, we presents the Value Compass Benchmarks, with three correspondingly designed modules. It (i) grounds the evaluation on motivationally distinct \textit{basic values to clarify LLMs' underlying values from a holistic view; (ii) applies a \textit{generative evolving evaluation framework with adaptive test items for evolving LLMs and direct value recognition from behaviors in realistic scenarios; (iii) propose a metric that quantifies LLMs alignment with a specific value as a weighted sum over multiple dimensions, with weights determined by pluralistic values.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。