构建首个面向AI研究创意生成的量化评估基准
AI Idea Bench 2025: AI Research Idea Generation Benchmark
- 基于3495篇论文构建可验证的创意生成评估数据集
- 从真实论文内容和通用知识双维度评估创意质量
- 解决大模型知识泄露与提示设计偏差问题,适合科研自动化研究者
大型语言模型(LLMs)在生成新想法方面取得显著进展,但现有评估方法忽视了模型的知识泄露、缺乏基于真实依据的开放性基准以及受提示设计限制的可行性分析。本文提出AI Idea Bench 2025,一个用于量化评估和比较LLMs在人工智能研究领域生成创意的框架。该框架包含3,495篇AI论文及其关联启发作品的综合性数据集,以及一套完整的评估方法。评估系统从两个维度衡量创意质量:与原始论文真实内容的一致性,以及基于通用参考材料的判断。该基准为评估和比较创意生成技术提供了重要资源,有助于推动科学发现的自动化。
原文摘要 · Abstract (English)
Large-scale Language Models (LLMs) have revolutionized human-AI interaction and achieved significant success in the generation of novel ideas. However, current assessments of idea generation overlook crucial factors such as knowledge leakage in LLMs, the absence of open-ended benchmarks with grounded truth, and the limited scope of feasibility analysis constrained by prompt design. These limitations hinder the potential of uncovering groundbreaking research ideas. In this paper, we present AI Idea Bench 2025, a framework designed to quantitatively evaluate and compare the ideas generated by LLMs within the domain of AI research from diverse perspectives. The framework comprises a comprehensive dataset of 3,495 AI papers and their associated inspired works, along with a robust evaluation methodology. This evaluation system gauges idea quality in two dimensions: alignment with the ground-truth content of the original papers and judgment based on general reference material. AI Idea Bench 2025's benchmarking system stands to be an invaluable resource for assessing and comparing idea-generation techniques, thereby facilitating the automation of scientific discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。