用大模型分析英国科研资助提案,识别新兴研究方向
Research Entity Extraction and Topic Detection from UKRI Grant Proposals
- 用Mistral模型提取提案中的研究实体并匹配开放主题库
- Mistral在主题分类上准确率达90.5%,远超其他方法
- 适合政策制定者追踪前沿科研趋势
本文报告了英国研究与创新署(UKRI)资助的元科学项目“追踪星辰与独角兽”的初步成果,对比了GPT-4o、Mistral及自研算法DSIT-Taxonomies三种基于大语言模型的方法,从科研资助提案中提取和分类研究实体。项目旨在识别新兴研究领域的早期信号,以指导公共投资。采用三阶段流程,利用Mistral进行主要实体抽取,并映射至OpenAlex Topics主题体系。在42份跨领域提案摘要上的评估显示,Mistral与GPT-4o生成的实体集质量相当且语义重叠度高,显著优于碎片化的DSIT-Taxonomies方法。关键结果:基于Mistral的主题分类准确率达90.5%,远超完整DSIT-Taxonomies流程的71.4%。结论表明,Mistral在大规模敏感资助数据处理中具备高性能、高效率与强安全性。
原文摘要 · Abstract (English)
This paper presents preliminary findings from a UKRI-funded Metascience project comparing three LLM-based approaches, GPT-4o, Mistral, and a bespoke algorithm, DSIT-Taxonomies, for extracting and classifying research entities from funding proposals. Our project "Tracking Stars and Unicorns" aims to identify early signals of emerging research areas to inform public investment. Our methodology employed a three-stage pipeline, leveraging Mistral for primary entity extraction and mapping against the OpenAlex Topics taxonomy. We evaluated our approach across 42 proposals' abstracts from different areas and observed that Mistral and GPT-4o produce comparable, high-quality entity sets with significant semantic overlap, outperforming the fragmented DSIT-Taxonomies approach. Crucially, the Mistral-based approach achieved superior topic classification accuracy (90.5%) compared to the full DSIT-Taxonomies pipeline (71.4%). We conclude that Mistral offers a high-performance, operationally efficient, and secure solution for large-scale analysis of sensitive grant data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。