现有大模型评测忽视个人偏好,新方法按用户个体差异排序模型。
Personalized Benchmarking: Evaluating LLMs by Individual Preferences

- 用ELO和布拉德利-特里模型为115名用户生成个性化模型排名
- 个体排名与整体平均排名相关性极低,仅43%用户有中等相关
- 用户主题兴趣和写作风格显著影响模型偏好,可据此预测偏好
随着大语言模型能力提升及其在真实任务中的部署,评估其与人类偏好的对齐程度成为重要挑战。当前基准测试将用户偏好取平均以计算整体评分,忽略了个体用户的偏好差异。由于用户在不同情境下偏好各异,我们提出应建立基于个体需求的个性化模型评测体系。通过为115名活跃的Chatbot Arena用户使用ELO评分和布拉德利-特里系数计算个性化模型排名,并分析用户提问特征(主题与写作风格)如何影响模型排名变化。结果表明,个体模型排名与总体排名差异显著:布拉德利-特里相关系数平均仅为ρ=0.04(57%的用户相关性接近零或为负),而ELO评分相关性中等(ρ=0.43)。通过主题建模与风格分析发现,用户在主题兴趣和沟通风格上存在显著异质性,进而影响其模型偏好。进一步证明,结合主题与风格特征的紧凑组合即可构成有效预测用户特定模型排名的特征空间。研究提供了量化证据:多数用户无法被聚合基准准确反映偏好,强调开发个性化评测体系的重要性。
原文摘要 · Abstract (English)
With the rise in capabilities of large language models (LLMs) and their deployment in real-world tasks, evaluating LLM alignment with human preferences has become an important challenge. Current benchmarks average preferences across all users to compute aggregate ratings, overlooking individual user preferences when establishing model rankings. Since users have varying preferences in different contexts, we call for personalized LLM benchmarks that rank models according to individual needs. We compute personalized model rankings using ELO ratings and Bradley-Terry coefficients for 115 active Chatbot Arena users and analyze how user query characteristics (topics and writing style) relate to LLM ranking variations. We demonstrate that individual rankings of LLM models diverge dramatically from aggregate LLM rankings, with Bradley-Terry correlations averaging only $ρ= 0.04$ (57\% of users show near-zero or negative correlation) and ELO ratings showing moderate correlation ($ρ= 0.43$). Through topic modeling and style analysis, we find users exhibit substantial heterogeneity in topical interests and communication styles, influencing their model preferences. We further show that a compact combination of topic and style features provides a useful feature space for predicting user-specific model rankings. Our results provide strong quantitative evidence that aggregate benchmarks fail to capture individual preferences for most users, and highlight the importance of developing personalized benchmarks that rank LLM models according to individual user preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。