arXiv:2505.01015cs.CLcs.AI2025-05ACL被引 19

用真实对话场景评估大模型价值观,更贴近实际使用。

Value Portrait: Assessing Language Models' Values through Psychometrically and Ecologically Valid Items

  • 基于真实用户-模型交互设计测试题
  • 44个模型显示重视利他、安全与自主
  • 发现模型对不同群体存在偏差,偏离人类数据

随着对更真实、符合人类价值观的大语言模型需求增加,评估其价值观的基准测试变得至关重要。然而,现有基准依赖人工或机器标注,易受价值偏见影响,且测试场景常脱离实际使用情境。为此,我们提出Value Portrait基准,具备两大特点:第一,题目源自真实用户-模型交互,提升评估结果与现实应用的相关性;第二,每道题由真人根据与自身思想的相似度打分,并计算评分与个体真实价值量表的相关性,确保高相关题目可作为可靠的价值评估工具。通过对44个大模型的评估,发现它们更倾向重视利他(Benevolence)、安全(Security)和自主(Self-Direction)价值观,而对传统(Tradition)、权力(Power)和成就(Achievement)关注度较低。此外分析揭示模型在看待不同人群时存在偏差,与真实人类数据不符。

原文摘要 · Abstract (English)

The importance of benchmarks for assessing the values of language models has been pronounced due to the growing need of more authentic, human-aligned responses. However, existing benchmarks rely on human or machine annotations that are vulnerable to value-related biases. Furthermore, the tested scenarios often diverge from real-world contexts in which models are commonly used to generate text and express values. To address these issues, we propose the Value Portrait benchmark, a reliable framework for evaluating LLMs' value orientations with two key characteristics. First, the benchmark consists of items that capture real-life user-LLM interactions, enhancing the relevance of assessment results to real-world LLM usage. Second, each item is rated by human subjects based on its similarity to their own thoughts, and correlations between these ratings and the subjects' actual value scores are derived. This psychometrically validated approach ensures that items strongly correlated with specific values serve as reliable items for assessing those values. Through evaluating 44 LLMs with our benchmark, we find that these models prioritize Benevolence, Security, and Self-Direction values while placing less emphasis on Tradition, Power, and Achievement values. Also, our analysis reveals biases in how LLMs perceive various demographic groups, deviating from real human data.

价值观评估大模型对齐心理测量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。