arXiv:2607.10511cs.CL2026-07

测试大模型评阅是否真懂论文,发现它常被表面流畅误导。

Articulate Intuition or Genuine Analysis? Benchmarking Epistemic Reliability in LLM-as-a-Judge Peer Reviews

论文配图:Articulate Intuition or Genuine Analysis? Benchmarking Epistemic Reliability in LLM-as-a-Judge Peer Reviews
图 1 · 摘自论文原文
  • 用心理学理论设计评分框架,量化评阅的思考深度与证据支撑
  • 大模型偏爱长文本和热门会议文章,但真实质量差异不大
  • 能区分深度分析与表面流畅,适合评估评阅可信度的研究者

当大模型判断一篇评阅‘分析性强’,而人类委员会认为另一篇‘质量高’时,它们是否在衡量同一标准?我们提出二者并不一致,且这种差异具有哲学意义。我们基于卡尼曼双系统理论构建评阅评价框架,发布Kahneman4Review基准数据集,包含3,563篇经评分的评阅,涵盖九个理论驱动的文本维度、八个偏差诊断项及连续推理质量得分。三项发现关乎可信度:决策层级与基于文本的真知质量代理指标无显著关联;公开展示的自主生成评阅获得更高原始分数,但长度与会议类型解释了大部分差距,且未进行论文配对;ICLR评阅诊断特征在2022–2023年发生转移,时间上与大模型普及重合,但无法确定因果。一项匹配式函数探测初步验证该框架可区分真正的批判性发现与仅表面流畅的表达。我们认为,可信的大模型评阅基准必须分离分析形式与认知功能,并提出具体设计方向。交互式演示可在https://huggingface.co/spaces/nuojohnchen/Kahneman4Review查看。

原文摘要 · Abstract (English)

When an LLM judge calls a peer review analytical and a human committee calls another review high quality, are they tracking the same thing? We argue they are not, and that the difference matters philosophically. We operationalise Kahneman's dual-process theory into a structured rubric for peer review and release Kahneman4Review, a benchmark of 3,563 rated reviews scored along nine theoretically motivated textual dimensions, eight bias diagnostics, and a continuous reasoning-quality score. Three findings bear on trustworthiness: decision tier is not detectably aligned with the rubric's text-grounded epistemic-quality proxy; public-showcase agentic reviews receive higher raw scores than pooled human reviews, but length and venue explain most of the gap and the samples are not paper-paired; and ICLR review-text diagnostics shift at the 2022--2023 transition, temporally coincident with widespread LLM availability but without identifying its cause. A matched function-probe pilot further shows that the rubric distinguishes textual probes designed to contrast genuine fault-finding with surface fluency. We argue that a trustworthy reliability benchmark for LLM judges must separate analytical form from epistemic function, and propose concrete design choices toward that goal. An interactive demo is available at https://huggingface.co/spaces/nuojohnchen/Kahneman4Review.

大模型评估评阅可信度认知科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。