首个自动化评估科研创意新颖性的基准,揭示大模型判断能力仍有短板。
Is this Idea Novel? An Automated Benchmark for Judgment of Research Ideas
- 构建1381个专家标注的科研创意数据集,配套九种评估指标。
- 大模型推理与人类理由高度相似,但新颖性判断准确率显著偏低。
- 适合研究AI辅助科研、可解释性评估或科学发现的学者参考。
判断科研创意的新颖性对推动科学发展至关重要,有助于识别未探索方向,并确保研究贡献真正拓展已有知识而非重复微小变体。然而,随着科学文献呈指数增长,通过人工文献综述判断创意新颖性既耗时又主观,难以规模化实施。为此,近期研究提出了自动化方法,但其评估方式缺乏统一标准,多依赖非标准化的人工评价,限制了大规模、可比性评估。本文提出RINoBench,首个面向科研创意新颖性判断的大规模评估基准,包含1,381个由专家生成并标注的科研创意,以及九种用于评估基于规则的新颖性评分和文本论证的指标。我们利用该基准评估多个前沿大语言模型(LLMs)在判断科研创意新颖性方面的能力。结果表明,尽管大模型生成的推理过程与人类理由高度一致,这种一致性并未转化为准确的新颖性判断——其结果与人类黄金标准存在显著偏差,即使在顶尖推理模型中亦然。数据与代码已公开:https://github.com/TimSchopf/RINoBench。
原文摘要 · Abstract (English)
Judging the novelty of research ideas is crucial for advancing science, enabling the identification of unexplored directions, and ensuring contributions meaningfully extend existing knowledge rather than reiterate minor variations. However, given the exponential growth of scientific literature, manually judging the novelty of research ideas through literature reviews is labor-intensive, subjective, and infeasible at scale. Therefore, recent efforts have proposed automated approaches for research idea novelty judgment. Yet, evaluation of these approaches remains largely inconsistent and is typically based on non-standardized human evaluations, hindering large-scale, comparable evaluations. To address this, we introduce RINoBench, the first comprehensive benchmark for large-scale evaluation of research idea novelty judgments. It comprises 1,381 research ideas derived from and judged by human experts as well as nine automated evaluation metrics designed to assess both rubric-based novelty scores and textual justifications of novelty judgments. Using this benchmark, we evaluate several state-of-the-art large language models (LLMs) on their ability to judge the novelty of research ideas. Our findings reveal that while LLM-generated reasoning closely mirrors human rationales, this alignment does not reliably translate into accurate novelty judgments, which diverge significantly from human gold standard judgments - even among leading reasoning-capable models. Data and code available at: https://github.com/TimSchopf/RINoBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。