构建科研创意生成评估基准,量化大模型的创新与可行性。
IdeaBench: Benchmarking Large Language Models for Research Idea Generation
- 将大模型模拟为领域研究员,基于真实论文上下文生成新创意。
- 用GPT-4o进行多维度评分,结合洞察分量化创意质量。
- 适合研究自动化科研发现、AI辅助创新的学者与工程师。
大语言模型(LLMs)已显著改变人机交互方式,在科学发现与假设生成等任务中取得顶尖成果。然而,缺乏系统化的评估框架来衡量其在科研创意生成方面的表现,成为制约理解与评估其生成能力的关键障碍。为此,我们提出IdeaBench,一个包含综合性数据集与评估框架的基准系统,用于标准化评测大模型生成科研创意的能力。数据集涵盖来自多个领域的重要论文标题与摘要及其引用文献。为模拟人类科研创意思维过程,我们将大模型定位为特定领域的研究人员,并将其置于与人类研究者相同的语境中,以最大化利用模型参数化知识,动态生成新研究方向。我们还设计了一套两阶段评估框架:首先使用GPT-4o根据用户指定的质量指标(如新颖性、可行性)对生成创意进行排序,实现可扩展的个性化评估;其次通过“洞察分”(Insight Score)计算相对排名,量化选定质量指标。该基准系统将为社区提供测量和比较不同大模型的能力工具,推动科研发现的自动化进程。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have transformed how people interact with artificial intelligence (AI) systems, achieving state-of-the-art results in various tasks, including scientific discovery and hypothesis generation. However, the lack of a comprehensive and systematic evaluation framework for generating research ideas using LLMs poses a significant obstacle to understanding and assessing their generative capabilities in scientific discovery. To address this gap, we propose IdeaBench, a benchmark system that includes a comprehensive dataset and an evaluation framework for standardizing the assessment of research idea generation using LLMs. Our dataset comprises titles and abstracts from a diverse range of influential papers, along with their referenced works. To emulate the human process of generating research ideas, we profile LLMs as domain-specific researchers and ground them in the same context considered by human researchers. This maximizes the utilization of the LLMs' parametric knowledge to dynamically generate new research ideas. We also introduce an evaluation framework for assessing the quality of generated research ideas. Our evaluation framework is a two-stage process: first, using GPT-4o to rank ideas based on user-specified quality indicators such as novelty and feasibility, enabling scalable personalization; and second, calculating relative ranking based "Insight Score" to quantify the chosen quality indicator. The proposed benchmark system will be a valuable asset for the community to measure and compare different LLMs, ultimately advancing the automation of the scientific discovery process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。