arXiv:2508.09463cs.CL2025-08被引 1

用用户偏好动态排名大模型,更贴合实际需求。

User-centric Subjective Leaderboard by Customizable Reward Modeling

  • 基于超1万条主观问答数据,构建可定制奖励模型
  • 40亿参数模型超越GPT-4.1与Gemini-2.5-pro表现
  • 适合关注个性化模型选择的开发者与研究者

现有大语言模型评估基准多聚焦于可验证任务,缺乏对用户实际需求的适配性。为此,我们提出首个以用户为中心的主观排行榜(USL),通过真实人类偏好数据实现跨场景动态排名。基于超过10,000条主观查询的调研发现,人类偏好存在显著多样性和矛盾性,制约了现有奖励模型效果。为此,我们引入可定制奖励模型(CRMs),仅需40亿参数即可超越GPT-4.1与Gemini-2.5-pro,在新主题和新标准下展现卓越泛化能力。由CRMs驱动的USL表现出与矛盾偏好强负相关性。

原文摘要 · Abstract (English)

Existing benchmarks for large language models (LLMs) predominantely focus on assessing their capabilities through verifiable tasks. Such objective and static benchmarks offer limited utility for practical LLM selection, making it difficult for users to find suitable models for their individual needs. To bridge this gap, we present the first User-Centric Subjective Leaderboard (USL), which provides a preference-driven, dynamic ranking of LLMs across diverse real-world scenarios. Our work is built upon a thorough investigation of real human preference data, involving more than 10K subjective queries. Our investigation reveals significant diversity and contradictions in human preferences, which limit the effectiveness of state-of-the-art reward models. To address this, we introduce Customizable Reward Models (CRMs). With only 4B parameters, our CRM surpasses the performance of leading models such as GPT-4.1 and Gemini-2.5-pro, showing exceptional generalization capabilities across new topics and criteria. The USL, powered by CRMs, exhibits strong negative correlations to contradictory preferences.

大模型评估用户偏好奖励建模动态排名

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。