arXiv:2510.06370cs.CL2025-10被引 2

评测大模型与评分模型对用户价值观和风格偏好的引导能力

EVALUESTEER: Measuring Reward Model Steerability Towards Values and Preferences

  • 构建16.6万组偏好数据,覆盖4类价值观与4种风格维度
  • 当前最优模型在完整用户画像下准确率不足75%
  • 揭示现有评分模型难以有效识别相关偏好信息

随着大型语言模型(LLMs)在全球部署,构建能够适应全球用户多样化偏好与价值观的系统变得至关重要。本文提出EVALUESTEER,一个基于心理学与人-模型交互文献的基准测试,用于衡量LLMs与评分模型(RMs)对用户价值与风格偏好配置的可引导性。为填补现有数据集无法支持可控评估的空白,我们合成生成了165,888组偏好对,系统性地在4个价值维度(传统、世俗理性、生存、自我表达)与4个风格维度(冗余度、可读性、自信度、亲和度)上进行变化。利用EVALUESTEER,评估在给定用户画像及一对价值/风格导向响应时,模型是否能选出符合用户偏好的输出。我们在十一种系统提示条件与六种偏好比较场景下,对六种开源与专有LLMs及RMs进行了评估。结果显示,当提供完整的用户价值与风格偏好画像时,最佳模型的准确率低于75%;而仅提供相关偏好时,准确率超过99%。这表明当前评分模型在识别与适配相关用户信息方面存在明显局限,并为开发可引导至多元人类价值观与偏好的评分模型提供了挑战性测试平台。

原文摘要 · Abstract (English)

As large language models (LLMs) are deployed globally, creating pluralistic systems that can accommodate the diverse preferences and values of users worldwide becomes essential. We introduce EVALUESTEER, a benchmark to measure LLMs' and reward models' (RMs) steerability towards users' value and stylistic preference profiles grounded in psychology and human-LLM interaction literature. To address the gap in existing datasets that do not support controlled evaluations of RM steering, we synthetically generated 165,888 preference pairs -- systematically varying pairs along 4 value dimensions (traditional, secular-rational, survival, and self-expression) and 4 style dimensions (verbosity, readability, confidence, and warmth). We use EVALUESTEER to evaluate whether, given a user profile and a pair of candidate value-laden and style-laden responses, LLMs and RMs are able to select the output that aligns with the user's preferences. We evaluate six open-source and proprietary LLMs and RMs under eleven systematic prompting conditions and six preference comparison scenarios. Notably, our results show that, when given the user's full profile of values and stylistic preferences, the best models achieve <75% accuracy at choosing the correct response, in contrast to >99% accuracy when only relevant style and value preferences are provided. EVALUESTEER thus highlights the limitations of current RMs at identifying and adapting to relevant user profile information, and provides a challenging testbed for developing RMs that can be steered towards diverse human values and preferences.

评分模型偏好引导价值观对齐评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。