arXiv:2604.23430cs.IRcs.AI2026-04

用大模型自动分类科学文献,提升研究信息管理效率

Automating Categorization of Scientific Texts with In-Context Learning and Prompt-Chaining in Large Language Models

论文配图:Automating Categorization of Scientific Texts with In-Context Learning and Prompt-Chaining in Large Language Models
图 1 · 摘自论文原文
  • 通过提示链技术让大模型理解多层次分类体系
  • 在领域和主题分类上准确率超现有模型,但三级标题仍不足50%
  • 适合需要自动化文献管理的研究机构与知识系统开发者

科学文献的爆炸式增长给知识发现和导航带来挑战。本研究系统评估了现成大语言模型(LLMs)在科学文本分类中的表现,采用分层的ORKG分类体系与FORC数据集作为基准。实验对比了上下文学习(ICL)和提示链(Prompt Chaining)策略,并探究温度参数对分类准确率的影响。结果表明,提示链在处理嵌套的ORKG分类结构时优于纯上下文学习,尤其在一级(领域)和二级(主题)分类上显著超越当前最优模型;然而在三级(具体主题)分类上,即使使用提示链,准确率也仅约50%,尚未达到理想水平。

原文摘要 · Abstract (English)

The relentless expansion of scientific literature presents significant challenges for navigation and knowledge discovery. Within Research Information Retrieval, established tasks such as text summarization and classification remain crucial for enabling researchers and practitioners to effectively navigate this vast landscape, so that efforts have increasingly been focused on developing advanced research information systems. These systems aim not only to provide standard keyword-based search functionalities but also to incorporate capabilities for automatic content categorization within knowledge-intensive organizations across academia and industry. This study systematically evaluates the performance of off-the-shelf Large Language Models (LLMs) in analyzing scientific texts according to a given classification scheme. We utilized the hierarchical ORKG taxonomy as a classification framework, employing the FORC dataset as ground truth. We investigated the effectiveness of advanced prompt engineering strategies, namely In-Context Learning (ICL) and Prompt Chaining, and experimentally explored the influence of the LLMs' temperature hyperparameter on classification accuracy. Our experiments demonstrate that Prompt Chaining yields superior classification accuracy compared to pure ICL, particularly when applied to the nested structure of the ORKG taxonomy. LLMs with prompt chaining outperform the state-of-the-art models for domain (1st level) prediction and show even better performance for subject (2nd level) prediction compared to the older BERT model. However, LLMs are not yet able to perform well in classifying the topic (3rd level) of research areas based on this specific hierarchical taxonomy, as they only reach about 50% accuracy even with prompt chaining.

文本分类大模型应用知识管理提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。