arXiv:2503.05712cs.CYcs.AI2025-03被引 9

用引用量和评分预测评估AI生成科研论文,发现标题摘要就能胜过大模型评审。

Automatic Evaluation Metrics for Artificially Generated Scientific Research

  • 基于标题摘要预测论文引用量,比预测评审分数更可靠。
  • 仅凭研究假说预测评分难度大,需完整论文才能更好判断。
  • 简单模型比LLM评审更准,但仍未达人类评审一致性水平。

基础模型在科研中应用日益广泛,但评估其生成的科研成果仍具挑战性。专家评审成本高,而以大语言模型作为代理评审者又不可靠。为此,我们研究了两种自动评估指标:引用量预测和评审分数预测。我们解析了OpenReview平台所有论文,为每篇提交增加引用数、参考文献及研究假设信息。结果表明,引用量预测比评审分数预测更可行;仅凭研究假设预测评分远难于使用完整论文。此外,仅基于标题和摘要的简单预测模型表现优于基于LLM的评审系统,但仍不及人类评审的一致性水平。

原文摘要 · Abstract (English)

Foundation models are increasingly used in scientific research, but evaluating AI-generated scientific work remains challenging. While expert reviews are costly, large language models (LLMs) as proxy reviewers have proven to be unreliable. To address this, we investigate two automatic evaluation metrics, specifically citation count prediction and review score prediction. We parse all papers of OpenReview and augment each submission with its citation count, reference, and research hypothesis. Our findings reveal that citation count prediction is more viable than review score prediction, and predicting scores is more difficult purely from the research hypothesis than from the full paper. Furthermore, we show that a simple prediction model based solely on title and abstract outperforms LLM-based reviewers, though it still falls short of human-level consistency.

AI评估引用预测论文评分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。