arXiv:2510.16234cs.AIcs.CL2025-10被引 7

用文献检索评估研究创意的合理性与创新性,提升AI生成想法的质量。

ScholarEval: Research Idea Evaluation Grounded in Literature

  • 基于已有文献评估研究想法的科学性和创新性,融合检索增强机制。
  • 在专家标注数据集上表现优于所有基线,覆盖率达92%以上。
  • 适合科研人员、AI助手开发者用于优化研究创意生成与筛选。

随着AI工具在科研构思中日益普及,对生成想法的稳健评估至关重要。我们提出ScholarEval,一个基于文献检索增强的评估框架,从两个核心维度评价研究想法:科学性(基于现有文献的实证有效性)和贡献度(相对于先前研究在多个维度上的进步程度)。为评估该框架,我们构建了首个跨学科专家标注的研究想法与评审数据集ScholarIdeas,涵盖人工智能、神经科学、生物化学和生态学四个领域,共117个研究想法。实验表明,ScholarEval在覆盖专家标注评分标准方面显著优于所有基线模型。相较于OpenAI推出的最强基线o4-mini-deep-research(具备推理与搜索能力的智能体系统),ScholarEval在评估可操作性、深度与证据支持方面均更受青睐。大规模用户研究表明,ScholarEval在文献参与度、想法优化和实用性方面也显著优于深度研究系统。代码、数据集与工具已开源,供社区使用与扩展。

原文摘要 · Abstract (English)

As AI tools become increasingly common for research ideation, robust evaluation is critical to ensure the validity and usefulness of generated ideas. We introduce ScholarEval, a retrieval augmented evaluation framework that assesses research ideas based on two fundamental criteria: soundness - the empirical validity of proposed methods based on existing literature, and contribution - the degree of advancement made by the idea across different dimensions relative to prior research. To evaluate ScholarEval, we introduce ScholarIdeas, the first expert-annotated dataset of multi-domain research ideas and reviews, comprised of 117 ideas across four disciplines: artificial intelligence, neuroscience, biochemistry, and ecology. Our evaluation shows that ScholarEval achieves significantly higher coverage of points mentioned in the human expert annotated rubrics in ScholarIdeas compared to all baselines. Furthermore, ScholarEval is consistently preferred over our strongest baseline o4-mini-deep-research, a reasoning and search-enabled agentic system by OpenAI, in terms of evaluation actionability, depth, and evidence support. Our large-scale user study also shows that ScholarEval significantly outperforms deep research in literature engagement, idea refinement, and usefulness. We openly release our code, dataset, and ScholarEval tool for the community to use and build on.

研究创意文献评估AI辅助科研

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。