用文本和引用匹配科研论文与专利,追踪知识转化效率。
Patent-publication pairs for the detection of knowledge transfer from research to industry: reducing ambiguities with word embeddings and references
- 结合作者、投资人、文本内容与引用共现识别论文专利对
- 基于医学术语向量化提升匹配准确率,五年内验证有效
- 公开完整流程,适合研究成果转化分析者使用
医疗研究的成效不仅体现在发表数量,也体现在经济可转化性。专利是研究成果产业化的体现,可作为科研到产业知识转移的代理指标。本研究旨在识别科研论文与专利的配对关系,以专利反映研究的经济影响。通过比对作者和资助机构名称进行匹配,并引入两项辅助特征:一是基于医学主题词表(MeSH)的技术术语向量化计算文本相似度,二是识别两类文献中的共同引用。在五年的示例期内评估了这两项支持特征的效果。此外,我们开发了一种统计方法,用于确定医学领域有效的专利分类。整个数据处理流程从原始文献数据到验证后的论文-专利对均开源可用。
原文摘要 · Abstract (English)
The performance of medical research can be viewed and evaluated not only from the perspective of publication output, but also from the perspective of economic exploitability. Patents can represent the exploitation of research results and thus the transfer of knowledge from research to industry. In this study, we set out to identify publication-patent pairs in order to use patents as a proxy for the economic impact of research. To identify these pairs, we matched scholarly publications and patents by comparing the names of authors and investors. To resolve the ambiguities that arise in this name-matching process, we expanded our approach with two additional filter features, one used to assess the similarity of text content, the other to identify common references in the two document types. To evaluate text similarity, we extracted and transformed technical terms from a medical ontology (MeSH) into numerical vectors using word embeddings. We then calculated the results of the two supporting features over an example five-year period. Furthermore, we developed a statistical procedure which can be used to determine valid patent classes for the domain of medicine. Our complete data processing pipeline is freely available, from the raw data of the two document types right through to the validated publication-patent pairs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。