评测大模型跨学科研究能力,发现其知识融合与创新潜力。
IDRBench: Understanding the Capability of Large Language Models on Interdisciplinary Research
- 构建跨学科研究评测框架IDRBench,包含三类任务。
- 在10个主流大模型上验证,揭示其知识整合能力差异。
- 适合研究AI辅助创新、跨学科智能的学者参考。
创新是人类文明发展的关键驱动力。随着知识体系不断扩展,跨学科领域间的知识融合——重大创新常在此处涌现——变得日益困难。近年来,机器学习模型(尤其是大语言模型,LLMs)在获取海量知识和推理方面表现优异,为跨学科发现带来新机遇。本研究旨在理解前沿大模型在跨学科研究(IDR)中整合多领域知识的能力。为此,我们提出IDRBench,首个系统性框架,包含三个评估任务:(1) 跨学科论文识别,(2) 跨学科思想融合,(3) 跨学科思想推荐。通过对10个主流大模型的全面分析,我们揭示了其行为模式,并建立基准线,为未来研究提供参考。据我们所知,IDRBench是首个对大模型跨学科能力进行综合性探究的工作。
原文摘要 · Abstract (English)
Innovation is a key driving force of human civilization. As the body of knowledge has grown considerably, bridging knowledge across different disciplines, where significant innovation often emerges, has become increasingly challenging. The recent advancements in machine learning models, particularly Large Language Models (LLMs), have provided effective access to extensive knowledge sources and shown impressive abilities in reasoning, rendering significant opportunities for interdisciplinary discovery. Our research aims to understand the capabilities of state-of-the-art LLMs in integrating knowledge from different fields for interdisciplinary research (IDR). To address this fundamental problem, we introduce IDRBench, a pioneering framework that includes both datasets and evaluation tasks: (1) IDR Paper Identification, (2) IDR Idea Integration, and (3) IDR Idea Recommendation. Our study on ten mainstream LLMs provides a comprehensive analysis of their behavior and establishes benchmarks and baselines for future research. To the best of our knowledge, IDRBench is the first to provide a comprehensive investigation of LLMs' IDR capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。