用文本分析筛选真正环保的专利,发现仅20%被标为绿色的专利是真绿。
Seeing Through Green: Text-Based Classification and the Firm's Returns from Green Patents
- 基于1240万专利训练神经网络,通过语义向量扩展绿色技术词典。
- 真实绿色专利占绿色专利总量20%,且被后续发明引用少1%。
- 拥有真实绿色专利的企业销售额、市场份额和生产率均显著提升。
本文提出一种自然语言处理方法,从官方支持文件中识别真正的绿色专利。我们以约1240万篇此前文献标记为绿色的专利为基础进行训练,构建简单神经网络,通过环境技术相关表达的向量表示扩展基础词典。测试结果显示,真正绿色的专利约占所有标记为绿色专利的20%。研究还发现不同技术类别间存在异质性,且真实绿色专利被后续发明引用次数低约1%。在论文第二部分,我们检验了专利与欧盟企业层面财务数据的关系,在控制反向因果后发现,持有至少一项真实绿色专利的企业,其销售额、市场占有率和生产率均有提升。若限定为高新颖性的真实绿色专利,则利润也更高。研究强调了利用文本分析实现更精细专利分类的重要性,对政策制定具有参考价值。
原文摘要 · Abstract (English)
This paper introduces Natural Language Processing for identifying ``true'' green patents from official supporting documents. We start our training on about 12.4 million patents that had been classified as green from previous literature. Thus, we train a simple neural network to enlarge a baseline dictionary through vector representations of expressions related to environmental technologies. After testing, we find that ``true'' green patents represent about 20\% of the total of patents classified as green from previous literature. We show heterogeneity by technological classes, and then check that `true' green patents are about 1\% less cited by following inventions. In the second part of the paper, we test the relationship between patenting and a dashboard of firm-level financial accounts in the European Union. After controlling for reverse causality, we show that holding at least one ``true'' green patent raises sales, market shares, and productivity. If we restrict the analysis to high-novelty ``true'' green patents, we find that they also yield higher profits. Our findings underscore the importance of using text analyses to gauge finer-grained patent classifications that are useful for policymaking in different domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。