不依赖生成时概率,用几何一致性预测模型评分与人工分歧
Predicting Disagreement with Human Raters in LLM-as-a-Judge Difficulty Assessment without Using Generation-Time Probability Signals

- 用独立嵌入空间分析评分集几何一致性判断分歧
- 在两个大模型上预测分歧的AUC优于基于概率的方法
- 适合需要降低人工标注成本的教育内容难度评估场景
使用大型语言模型(LLMs)自动生成教育材料日益普遍,但为其分配难度等级仍需大量人工投入。因此,以LLM为评判者的方法受到关注,但其与人工评分的分歧仍是主要挑战。本文提出一种预测哪些LLM生成的难度评分可能与人工评分不一致的方法,以便将此类情况送交重新评定。不同于以往依赖生成时概率信号的方法(该信号需在评分生成阶段收集,且跨模型难以比较),本方法利用难度为有序尺度的特性,采用如ModernBERT等独立嵌入空间,通过评分集的几何一致性识别潜在分歧样本。在基于CEFR的英文句子难度评估任务上,使用GPT-OSS-120B和Qwen3-235B-A22B进行实验,结果表明,所提方法在预测分歧方面的AUC高于基于概率的基线方法。
原文摘要 · Abstract (English)
Automatic generation of educational materials using large language models (LLMs) is becoming increasingly common, but assigning difficulty levels to such materials still requires substantial human effort. LLM-as-a-Judge has therefore attracted attention, yet disagreement with human raters remains a major challenge. We propose a method for predicting which LLM-generated difficulty ratings are likely to disagree with human raters, so that such cases can be sent for re-rating. Unlike prior approaches, our method does not rely on generation-time probability signals, which must be collected during rating generation and are often difficult to compare across LLMs. Instead, exploiting the fact that difficulty is an ordinal scale, we use a separate embedding space, such as ModernBERT, and identify disagreement candidates based on the geometric consistency of the rating set. Experiments on English CEFR-based sentence difficulty assessment with GPT-OSS-120B and Qwen3-235B-A22B showed that the proposed method achieved higher AUC for predicting disagreement with human raters than probability-based baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。