arXiv:2605.04972cs.CL2026-05

专家评估难对齐,因判断风格差异大且依赖隐性标准。

Why Expert Alignment Is Hard: Evidence from Subjective Evaluation

  • 通过专家评估与问卷分析不同信息形式对对齐的影响。
  • 专家间判断差异大,部分维度对齐难度高,尤其需外部知识时。
  • 明确规则未必提升对齐效果,小样本修改虽有用但不稳定。

将大型语言模型与专家判断对齐在主观评价任务中尤为困难,因专家可能意见不一、依赖隐性标准并随时间改变判断。本文通过专家评估与后续问卷,研究不同形式的专家信息如何影响对齐,并揭示主观判断的本质。研究发现四个一致模式:第一,专家间对齐难度差异显著,表明其评估风格与模型先验行为的距离各异;第二,明确标准与推理并不总能提升对齐,说明专家判断未被言语化规则完全捕捉;第三,编辑效果对示例数量和身份均敏感,少量编辑可带来有效但不稳定的改善;第四,对齐难度因评价维度而异:更直接基于内容的维度较易对齐,而依赖外部知识或价值判断的维度仍困难。综合来看,专家对齐之难不仅源于模型局限,更因主观评价本身具有异质性、部分隐性、维度依赖及时间不稳定性。

原文摘要 · Abstract (English)

Aligning large language models with expert judgment is especially difficult in subjective evaluation tasks, where experts may disagree, rely on tacit criteria, and change their judgments over time. In this paper, we study expert alignment as a way to understand this difficulty. Using expert evaluations and follow-up questionnaires, we examine how different forms of expert information affect alignment and what this reveals about subjective judgment. Our findings show four consistent patterns. First, alignment difficulty varies substantially across experts, suggesting that expert evaluation styles differ widely in their distance from a model's prior behavior. Second, explicit criteria and reasoning do not always improve alignment, indicating that expert judgment is not fully captured by verbalized rules. Third, editing is sensitive to both the number and the identity of examples, with small numbers of edits providing useful but unstable gains. Fourth, alignment difficulty differs across evaluation dimensions: dimensions grounded more directly in proposal content are easier to align, while dimensions requiring external knowledge or value-based judgment remain harder. Taken together, these results suggest that expert alignment is difficult not only because of model limitations, but also because subjective evaluation is inherently heterogeneous, partly tacit, dimension-dependent, and temporally unstable.

模型对齐主观评估专家判断人类反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。