用大模型从专利中提取关键概念,自动构建可信的可持续发展目标标签数据。
From scratch to silver: Creating trustworthy training data for patent-SDG classification using Large Language Models
- 基于专利引用文献中的目标标签,结合大模型提取功能、方案和应用三类结构化概念。
- 通过排名检索融合跨领域相似度,训练出可识别新关联的软标签数据集。
- 适用于需要大规模专利与可持续发展目标匹配的研究者或政策分析人员。
将专利与联合国可持续发展目标(SDGs)进行分类,对追踪创新如何应对全球挑战至关重要。然而,缺乏大规模标注数据限制了监督学习的应用。现有方法如关键词搜索、迁移学习和引文启发式规则在可扩展性和泛化性上存在不足。本文将专利-SDG分类视为弱监督问题,利用专利引用已标记为SDG的非专利文献(NPL引用)作为噪声初始信号。针对其稀疏性和噪声,我们设计了一个复合标注函数(LF),使用大语言模型(LLMs)基于专利本体从专利和SDG论文中提取结构化概念:功能、解决方案与应用。通过基于排名的检索方法计算跨领域相似度并融合。该标注函数采用自定义的仅正例损失进行校准,与已知的NPL-SDG关联对齐,同时不惩罚发现新关联。最终生成一个银标准的软多标签数据集,实现专利到SDG的有效映射,并支持多标签回归模型训练。我们通过两种互补策略验证方法:(1) 内部验证:与保留的基于NPL的标签对比,优于多个基线包括基于Transformer的模型和零样本大模型;(2) 外部验证:利用专利引用网络、共同发明人和共同申请人图的模块度,结果表明我们的标签在主题、认知与组织一致性上优于传统技术分类。证明弱监督与语义对齐可在大规模下提升SDG分类效果。
原文摘要 · Abstract (English)
Classifying patents by their relevance to the UN Sustainable Development Goals (SDGs) is crucial for tracking how innovation addresses global challenges. However, the absence of a large, labeled dataset limits the use of supervised learning. Existing methods, such as keyword searches, transfer learning, and citation-based heuristics, lack scalability and generalizability. This paper frames patent-to-SDG classification as a weak supervision problem, using citations from patents to SDG-tagged scientific publications (NPL citations) as a noisy initial signal. To address its sparsity and noise, we develop a composite labeling function (LF) that uses large language models (LLMs) to extract structured concepts, namely functions, solutions, and applications, from patents and SDG papers based on a patent ontology. Cross-domain similarity scores are computed and combined using a rank-based retrieval approach. The LF is calibrated via a custom positive-only loss that aligns with known NPL-SDG links without penalizing discovery of new SDG associations. The result is a silver-standard, soft multi-label dataset mapping patents to SDGs, enabling the training of effective multi-label regression models. We validate our approach through two complementary strategies: (1) internal validation against held-out NPL-based labels, where our method outperforms several baselines including transformer-based models, and zero-shot LLM; and (2) external validation using network modularity in patent citation, co-inventor, and co-applicant graphs, where our labels reveal greater thematic, cognitive, and organizational coherence than traditional technological classifications. These results show that weak supervision and semantic alignment can enhance SDG classification at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。