arXiv:2508.13182cs.DLcs.AI2025-08

用大模型模拟专家直觉,自动分类科研摘要并发现资助趋势

Using Artificial Intuition in Distinct, Minimalist Classification of Scientific Abstracts for Management of Technology Portfolios

  • 用大模型生成元数据,模仿专家直觉进行抽象分类
  • 在中美双数据集上验证,实现可区分的标签体系
  • 适合科技战略管理、技术侦察等需要宏观洞察的场景

科研摘要分类对战略决策有帮助,但因文本稀疏难自动化。现有方法依赖元数据提升性能,常需半监督设置,且标签易重叠、缺乏区分性。专家则能轻松标注与排序。本文提出一种称作「人工直觉」的流程,利用大语言模型(LLM)生成元数据,复现专家思路。基于美国国家科学基金会公开摘要构建标签集,并在中文国家自然科学基金摘要上测试,分析资助趋势。结果表明该方法适用于科研组合管理、技术探查等战略活动。

原文摘要 · Abstract (English)

Classification of scientific abstracts is useful for strategic activities but challenging to automate because the sparse text provides few contextual clues. Metadata associated with the scientific publication can be used to improve performance but still often requires a semi-supervised setting. Moreover, such schemes may generate labels that lack distinction -- namely, they overlap and thus do not uniquely define the abstract. In contrast, experts label and sort these texts with ease. Here we describe an application of a process we call artificial intuition to replicate the expert's approach, using a Large Language Model (LLM) to generate metadata. We use publicly available abstracts from the United States National Science Foundation to create a set of labels, and then we test this on a set of abstracts from the Chinese National Natural Science Foundation to examine funding trends. We demonstrate the feasibility of this method for research portfolio management, technology scouting, and other strategic activities.

文本分类大模型应用科技管理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。