评测大模型在学术引用验证中的表现,发现便宜模型也能胜任。
Do You Need a Frontier Model as a Citation Verifier? Benchmarking Rubric LLMs for Deep-Research Source Attribution
- 用多个现成大模型评估引用质量,按相关性和事实支持双维度打分
- 378个难题案例中,最小模型在相关性上达F1 0.908,其他模型表现相近
- 不同模型存在偏差差异,仅看准确率会掩盖方向性错误,影响强化学习训练
深度研究系统要求生成内容必须有可靠引文支撑。本文针对引文质量这一结构化评分任务,评估了来自三个模型家族的8个现成大模型作为评判者的表现。在包含1248个评分决策的对抗性长文本基准上,所有判断均经人工复核,其中378个为因模型分歧而判定的难题案例。结果显示,廉价模型在源相关性维度表现优异(GPT-5-mini F1达0.908,κ=0.636),而在事实支持维度各模型无显著差异。尽管总体F1相近,但模型间在通过率漂移、假阳性率和假阴性率上差异明显。单一数值的F1会掩盖这种方向性偏差,而这正是下游强化学习可能强化的问题。因此,使用引文评分作为奖励信号前,必须对评判模型进行校准,且无需最昂贵的模型。
原文摘要 · Abstract (English)
Reinforcement learning increasingly relies on an LLM judge to score each rubric criterion, and that judge acts as the reward model during training. Before such a signal can be trusted, we need to know how capable the judge must be and how biased it is. We study this calibration question for citation quality in deep-research systems, where a search-grounded LLM must support each claim it writes with a cited source. Citation quality is a structured rubric task in which each attribution-citation pair is judged along two dimensions that require an LLM, source relevance and factual support. On an adversarial long-form benchmark, we score 8 off-the-shelf LLM judges from 3 model families against gold labels over 1,248 rubric decisions, all of which were human-reviewed and 378 of which were hard cases adjudicated from judge disagreements. Cheaper judges remain competitive across both dimensions, with GPT-5-mini attaining the strongest source-relevance pass-class F1 at 0.908 ($κ$=0.636), while on factual support the judges are statistically indistinguishable (overlapping confidence intervals), so no single model dominates. At comparable F1, the judges still differ substantially in pass-rate drift, false positive rate, and false negative rate. Scalar F1 obscures this directional bias, yet it is exactly what a downstream reinforcement learning loop would reinforce. Calibrating the judge is therefore a prerequisite for using citation rubrics as reward signals, and our results show that this calibration does not require the most expensive available model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。