用专家对决方式评估文献综述生成质量,发现现有模型表现仍差。
LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform

- 设计对抗式评审平台,让专家匿名对比文献综述草案。
- 最强模型仅23%胜率,而智能体模型表现远超基础大模型。
- 提出新评测工具LitJudge,更贴近专家判断,适合研究者使用。
文献综述是科学进步的关键,但自动生成的综述难以评估,因其价值依赖专家判断而非简单的参考重叠。本文提出LitReview Arena,一种专为文献综述设计的对抗式评审平台:领域专家(具AI写作经验)在自身专业领域内匿名比较生成稿,按五项特定标准评分。共收集约3000条专家判断,每条包含五维评分。结果显示,当前最强模型在整体实用性上仅23.0%胜过人类稿,而智能体类LLM(如Sonar Deep Research)相比基线模型提升超60%。此外,现有LLM作为裁判的方法与专家意见显著不符(斯皮尔曼相关系数0.467),尤其在结构与研究建议等合成性标准上。基于收集偏好数据,我们构建了专家校准的评测器LitJudge,相关性提升至0.78,接近专家间一致性。代码与数据已公开于https://github.com/VanellopeAsher/LitReview-Arena。
原文摘要 · Abstract (English)
Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspects of research utility depend on expert judgment rather than reference-overlap metrics. We introduce LitReview Arena, a battle-style evaluation platform with a structured protocol tailored to literature review quality: domain experts with AI paper-writing experience compare anonymized drafts, are matched to topics within their expertise, and provide dimension-wise outcomes over five literature-review-specific criteria. From this protocol, we collect approximately 3k expert judgments, each containing five dimension-wise outcomes, and show that even the strongest current systems win only 23.0% of decisive matches against human drafts on overall utility, while agentic LLMs such as Sonar Deep Research substantially outperform base language models by over 60%. We further find that existing LLM-as-a-judge methods are substantially misaligned with human experts (Spearman's rho=0.467), especially on synthesis-heavy criteria such as paper structure and research suggestions. Using the collected preference data, we provide an expert-calibrated evaluator, LitJudge, which improves alignment to Spearman's rho=0.78, comparable to inter-expert consistency; code and data are publicly available at https://github.com/VanellopeAsher/LitReview-Arena.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。