用深度学习模型追踪科技文献中的知识流动,识别专利与论文间的语义关联。
Tracing the Flow of Knowledge From Science to Technology Using Deep Learning
- 基于SPECTER2微调出Pat-SPECTER模型,统一处理专利与论文语义相似性。
- 在预测专利-论文引用关系上表现最佳,准确率显著优于其他8个模型。
- 揭示美国专利引用的论文语义相似度较低,可能与披露义务有关,适合政策研究者使用。
我们开发了一种适用于专利与科学出版物的语义相似性模型。在类似竞速的评估中,将八个语言(相似性)模型用于预测可信的专利-论文引用关系。结果表明,我们的Pat-SPECTER模型表现最佳,该模型是SPECTER2在专利数据上微调得到。在两个实际应用场景(分离专利-论文对、预测专利-论文对)中,展示了Pat-SPECTER的能力。最后,我们验证了假设:美国专利引用的论文语义相似度低于其他主要司法管辖区,这可能源于披露义务。该模型对学术界和实践者均开放。
原文摘要 · Abstract (English)
We develop a language similarity model suitable for working with patents and scientific publications at the same time. In a horse race-style evaluation, we subject eight language (similarity) models to predict credible Patent-Paper Citations. We find that our Pat-SPECTER model performs best, which is the SPECTER2 model fine-tuned on patents. In two real-world scenarios (separating patent-paper-pairs and predicting patent-paper-pairs) we demonstrate the capabilities of the Pat-SPECTER. We finally test the hypothesis that US patents cite papers that are semantically less similar than in other large jurisdictions, which we posit is because of the duty of candor. The model is open for the academic community and practitioners alike.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。