用大模型自动生成论文复现评估标准,效果接近人工水平
Can LLMs Write Reliable Rubrics? A Meta-Evaluation for Experiment Reproduction

- 将评估标准转为清单格式,让大模型生成可执行的评分条目
- 生成标准与人工标准对齐度接近人类水平,但细节过细且倾向高分
- 适合想快速构建可复现性评测框架的研究者使用
基于评分表的评估是衡量大模型研究代理开放输出的有前景方法,尤其在论文复现任务中,直接对比论文与代码仓库易产生幻觉。然而,定制化论文评分表需大量专家投入,限制了PaperBench等基准的可扩展性。本文首次系统性地对大模型生成的评分表进行元评估。我们将评分表重构为清单形式,在两个主干模型上测试四种生成设置。通过语义相似性(内在)和评分一致性(外在)双重评估,结果表明增强设置显著提升下游评估对齐度,最强方案接近人工基线;而内在改进较有限。进一步分析显示,大模型生成的评分表往往过于细致、倾向高分,且对论文领域适应性差,揭示其优势与局限。
原文摘要 · Abstract (English)
Rubric-based evaluation is a promising approach for assessing open-ended outputs from LLM-based research agents, particularly in paper reproduction, where direct paper-to-repository comparison is prone to hallucination. However, constructing paper-specific rubrics requires substantial expert effort, limiting the scalability of benchmarks such as PaperBench. In this work, we present, to our knowledge, the first systematic meta-evaluation of LLM-generated rubrics for paper reproduction. We reformulate rubrics into a checklist-style format and evaluate four generation settings across two backbone models. We meta-evaluate generated rubrics intrinsically by semantic similarity and extrinsically by score alignment with ground-truth rubrics. Our results show that the augmented settings substantially improves downstream evaluation alignment, with the strongest setting approaching the human baseline, while intrinsic gains are more modest. Further analyses reveal that LLM-generated rubrics are often overly fine-grained, biased toward high scores, and less adaptive to paper domains, highlighting both the affordances and limitations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。