arXiv:2605.04875cs.CLcs.AI2026-05

用专利语言预测未来技术组合,提前数十年发现创新线索。

Anticipating Innovation Using Large Language Models

论文配图:Anticipating Innovation Using Large Language Models
图 1 · 摘自论文原文
  • 将专利分类码当作词汇,用Transformer模型学习技术语言。
  • 技术描述的集体语义趋同度可提前几十年预测首次技术组合。
  • 适合科技政策、研发规划和早期创新追踪的从业者使用。

预测创新——即新科技组合的出现——是科学与政策的核心挑战。我们发现,未来的技术组合会在专利的集体语言中留下早期痕迹,其预测信号甚至可在数十年前就被识别。这种信号并非来自单一发明人,而是数千项专利中技术描述方式的集体变化所致。为此,我们提出TechToken,一个基于Transformer的模型,将国际专利分类码(IPC)视为词汇,通过微调嵌入这些代码以学习技术语言。我们定义代码嵌入之间的上下文相似度作为语言趋同度量,结果表明该指标能准确预测首次技术组合。此外,TechToken在多项专利相关任务中优于现有先进模型,显著提升技术表示质量。

原文摘要 · Abstract (English)

Forecasting innovation, intended as the emergence of new technological combinations, is a fundamental challenge for science and policy. We show that forthcoming combinations leave an early trace in the collective language of patents, with predictive signals detectable even decades in advance. We show that signal is not attributable to any single inventor, but emerges as a collective shift in how technologies are described across thousands of patents. To this end, we introduce TechToken, a transformer-based model that treats technologies, classified by International Patent Classification codes, as words in its vocabulary, learning the language of technologies by embedding these codes during fine-tuning. We define context similarity between code embeddings as a measure of linguistic convergence and show that it accurately predicts first technological combinations. TechToken also improves general representation quality, outperforming state-of-the-art models across different patent-related tasks.

技术创新大模型专利分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。