arXiv:2412.11948cs.AI2024-12NAACL被引 66

用80亿参数模型生成更真实、批判性强的学术论文评审。

OpenReviewer: A Specialized Large Language Model for Generating Critical Scientific Paper Reviews

  • 基于7.9万篇顶会专家评审微调,专精于学术评审
  • 对400篇测试论文生成的评审更接近人类评分分布
  • 适合想快速改进论文的作者使用

我们提出OpenReviewer,一个开源系统,用于生成机器学习与人工智能会议论文的高质量同行评审。核心是经过7.9万份顶级会议专家评审微调的80亿参数语言模型Llama-OpenReviewer-8B。给定论文PDF和评审模板,该系统可提取全文内容(包括公式和表格),并依据会议规范生成结构化评审。在400篇测试论文上的评估显示,相比通用大模型GPT-4和Claude-3.5,OpenReviewer生成的评审更具批判性且更贴近真实。其他模型常给出过于积极的评价,而OpenReviewer的建议与人类评审评分分布高度一致。该系统为作者提供快速、建设性反馈以改进稿件,但不取代人工评审。OpenReviewer已开放在线演示和源代码。

原文摘要 · Abstract (English)

We present OpenReviewer, an open-source system for generating high-quality peer reviews of machine learning and AI conference papers. At its core is Llama-OpenReviewer-8B, an 8B parameter language model specifically fine-tuned on 79,000 expert reviews from top conferences. Given a PDF paper submission and review template as input, OpenReviewer extracts the full text, including technical content like equations and tables, and generates a structured review following conference-specific guidelines. Our evaluation on 400 test papers shows that OpenReviewer produces considerably more critical and realistic reviews compared to general-purpose LLMs like GPT-4 and Claude-3.5. While other LLMs tend toward overly positive assessments, OpenReviewer's recommendations closely match the distribution of human reviewer ratings. The system provides authors with rapid, constructive feedback to improve their manuscripts before submission, though it is not intended to replace human peer review. OpenReviewer is available as an online demo and open-source tool.

论文评审LLM应用AI工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。