arXiv:2506.22026cs.IRcs.AI2025-06被引 17

用大模型自动评估科学创意的新颖性,比现有方法高13%准确率。

Literature-Grounded Novelty Assessment of Scientific Ideas

  • 先关键词检索,再嵌入过滤+分面重排,层层筛选相关文献。
  • 在多个数据集上实现约13%的评价一致率提升。
  • 适合需要高效筛选创新点的研究者和AI辅助研发团队。

自动化科学创意生成系统进展显著,但创意新颖性的自动评估仍是一个关键且研究不足的挑战。人工通过文献回顾评估新颖性耗时费力,易受主观影响,难以规模化。为此,我们提出Idea Novelty Checker,一个基于大模型的检索增强生成(RAG)框架,采用两阶段检索-重排策略:首先通过关键词和片段检索获取广泛相关论文,再经嵌入过滤和分面式大模型重排进行精炼。系统引入专家标注样本,引导模型对比文献并生成基于文献的推理。大量实验表明,该新颖性检查器相比现有方法达成约13%更高的评价一致性。消融实验进一步验证了分面重排模块在识别最相关文献中的关键作用。

原文摘要 · Abstract (English)

Automated scientific idea generation systems have made remarkable progress, yet the automatic evaluation of idea novelty remains a critical and underexplored challenge. Manual evaluation of novelty through literature review is labor-intensive, prone to error due to subjectivity, and impractical at scale. To address these issues, we propose the Idea Novelty Checker, an LLM-based retrieval-augmented generation (RAG) framework that leverages a two-stage retrieve-then-rerank approach. The Idea Novelty Checker first collects a broad set of relevant papers using keyword and snippet-based retrieval, then refines this collection through embedding-based filtering followed by facet-based LLM re-ranking. It incorporates expert-labeled examples to guide the system in comparing papers for novelty evaluation and in generating literature-grounded reasoning. Our extensive experiments demonstrate that our novelty checker achieves approximately 13% higher agreement than existing approaches. Ablation studies further showcases the importance of the facet-based re-ranker in identifying the most relevant literature for novelty evaluation.

科学创新大模型新颖性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。