arXiv:2606.03189cs.CL2026-06ACL

让AI judge更懂人,基于真实对话偏好定制评估框架

SenseJudge: Human-Centric Preference-Driven Judgment Framework

论文配图:SenseJudge: Human-Centric Preference-Driven Judgment Framework
图 1 · 摘自论文原文
  • 用真实多轮对话构建人类偏好数据,驱动可定制的AI评判
  • 在个性化判官任务中超越现有方法,模型排名更贴近真人判断
  • 适合需要人性化评估的对话系统研发与评测场景

大型语言模型作为各类场景下的评判者,如评估模型回复,正成为主流范式。然而,现有评判方法多依赖固定偏好数据训练的评判模型,难以反映多样化的用户偏好,也难适应真实人机对话环境。为此,我们提出SenseJudge——一个以人类偏好驱动的可定制化评判框架,并构建了基于真实多轮交互的SenseBench基准测试集。将自动评判框架与基准应用于两个任务:(1)大模型作为个性化评判者;(2)模型排序。大量实验表明,SenseJudge在个性化判官任务中优于其他方法,且模型排序结果与真实人类感知高度一致。此外,我们还分析了位置偏差与一致性问题,并通过消融实验验证了框架的鲁棒性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) as judges across various scenarios such as assessing model responses is becoming an increasingly accepted paradigm. However, existing judgment approaches often rely on trained judgers using fixed preference data, which tend to overlook diverse user preferences and struggle to adapt to real-world human-AI dialogue scenarios. To address these limitations, we propose SenseJudge, a customizable judgment framework driven by human preferences and SenseBench, a diverse and challenging instruction-following benchmark derived from real-world multi-turn interactions. We applied the automatic judgment framework and benchmark to two tasks: (1) LLMs as personalized judges, and (2) model ranking. We conducted extensive experiments, and the results demonstrate that the SenseJudge framework surpasses other judgment methods and models in the LLMs-as-personalized-judges task and achieves model ranking that aligns with real human sense. Additionally, we conducted analyses on position bias and consistency, alongside ablation studies, which affirmed the robustness of SenseJudge.

AI评估人类偏好对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。