arXiv:2603.03319cs.CLcs.AI2026-03被引 1

自动发现大模型评价偏好中的隐藏概念,揭示其与人类不同的判断倾向。

Automated Concept Discovery for LLM-as-a-Judge Preference Analysis

  • 用稀疏自编码器从嵌入层提取可解释的偏好特征
  • 发现模型更倾向具体、共情表达,反感主动法律建议
  • 无需预设偏见类别,适合研究模型评价机制的人

大型语言模型(LLMs)被广泛用作模型输出的可扩展评估者,但其偏好判断存在系统性偏差,且常与人类评估不一致。以往研究多聚焦于少数预设偏见,未能自动发现未知的偏好驱动因素。本文通过分析多种嵌入层概念提取方法,比较其可解释性与预测能力,发现基于稀疏自编码器的方法能提取出更具可解释性的偏好特征,同时在预测模型决策方面表现不俗。基于超过27,000对来自多个真人偏好数据集的响应及三组LLM的判断结果,我们验证了现有结论(如模型比人类更频繁拒绝敏感请求),并发现了新趋势:在通用和领域特定数据中,模型偏好强调情境中的具体性和共情性;在学术建议中偏好细节与正式表达;对鼓励采取积极行动(如报警、起诉)的法律建议存在负面倾向。结果表明,自动化概念发现可无需预设偏见分类体系,实现对LLM评价偏好的系统性分析。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly used as scalable evaluators of model outputs, but their preference judgments exhibit systematic biases and can diverge from human evaluations. Prior work on LLM-as-a-judge has largely focused on a small, predefined set of hypothesized biases, leaving open the problem of automatically discovering unknown drivers of LLM preferences. We address this gap by studying several embedding-level concept extraction methods for analyzing LLM judge behavior. We compare these methods in terms of interpretability and predictiveness, finding that sparse autoencoder-based approaches recover substantially more interpretable preference features than alternatives while remaining competitive in predicting LLM decisions. Using over 27k paired responses from multiple human preference datasets and judgments from three LLMs, we analyze LLM judgments and compare them to those of human annotators. Our method both validates existing results, such as the tendency for LLMs to prefer refusal of sensitive requests at higher rates than humans, and uncovers new trends across both general and domain-specific datasets, including biases toward responses that emphasize concreteness and empathy in approaching new situations, toward detail and formality in academic advice, and against legal guidance that promotes active steps like calling police and filing lawsuits. Our results show that automated concept discovery enables systematic analysis of LLM judge preferences without predefined bias taxonomies.

大模型评估偏好分析概念发现可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。