用人类参考标准评估75个大模型的价值优先级,发现模型常夸大价值差异。
Evaluating AI Alignment in LLMs: Output Analysis of Value Priorities Across 75 Models with Human Benchmarking
- 将模型输出视为行为数据,对比其价值排序与人类基准
- 75个模型中多数排序正确但差距被系统性放大
- 该方法适合关键场景前的对齐审计,不依赖模型大小或能力
大型语言模型(LLMs)在人机交互中广泛应用,但现有能力与安全评测难以揭示其表达的价值优先级及其与人类判断的对应关系。本研究通过三个实验引入一种基于输出的评估方法:将模型生成文本视为行为数据,与人类参考进行比对。研究1通过归纳质性分析提炼出六类理想AI表现主题:性能、适应能力、社会福祉、伦理与责任、关系整合、自主性。研究2表明模型输出具有高度稳定性,跨模型价值优先结构趋于一致,具备可比性。研究3将75个主流模型与376名人类受访者进行对比,采用包含优先级顺序和差值校准的“轮廓保真度”指标。尽管多数模型保持了与人类一致的价值排序,但部分模型系统性夸大了各价值间的差异,说明模型可能在传统基准上看似对齐,实则偏离人类价值校准。轮廓保真度在不同模型间差异显著,且不随模型规模、发布时间或能力层级单调变化。无论是模型还是人类,均趋向弱化自主性,引发对日益自主型AI发展的反思。该研究提出的六类主题及轮廓评估方法,为部署前的关键场景对齐审计提供了可扩展工具。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used in human-AI interaction research and practice, yet existing capability and safety benchmarks reveal little about the value priorities these systems express or how those priorities correspond to human judgements. Across three studies, we introduce an output-based approach to evaluating one facet of AI alignment by treating LLM-generated text as behavioural data and comparing expressed value-priority profiles with a human reference. Study 1 used inductive qualitative analysis to derive six themes of optimal AI functioning, namely Performance, Adaptive Capacity, Social Good, Ethics and Responsibility, Relational Integration, and Agency. Study 2 showed that LLM outputs were highly stable within models and converged on a common value-priority structure across models, indicating reliable and comparable value profiles. Study 3 benchmarked 75 contemporary LLMs against 376 human respondents using a profile-fidelity metric capturing both the relative ordering of priorities and the calibration of between-priority differences. Although most models reproduced the human ordering of values, some systematically exaggerated the differences between them, showing that models can appear aligned on conventional benchmarks while still diverging from human value calibration. Profile fidelity varied substantially across models and did not consistently scale with size, recency, or capability tier. Both LLMs and humans converged on a deprioritisation of Agency, raising important questions about the development of increasingly agentic AI systems. For research and applied use, the six themes and profile-based metric provide a scalable method for auditing LLM value profiles before deployment in contexts where alignment with human priorities is critical.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。