arXiv:2607.19178cs.CL2026-07

自动提取7.6万篇能源模型论文数据,构建可分析的公开数据库

Automated Extraction of Techno-Economic Data from 76,000 Energy System Studies

论文配图:Automated Extraction of Techno-Economic Data from 76,000 Energy System Studies
图 1 · 摘自论文原文
  • 用自动化方法从7.6万篇论文中提取定量数据
  • 建成包含320万条数据点的结构化数据库
  • 适合能源政策研究者和建模人员使用

能源系统模型指导重大社会决策,但其可信度依赖于难以溯源和审计的量化假设。元分析可提升透明度与建模规范性,但文献数量激增使人工信息提取日益不可行,导致数据库更新滞后、研究重复。本文展示了对2010年以来7.6万篇能源系统研究的高精度自动化数据提取。我们构建了包含320万条结构化定量数据点及2000万条关联元数据的数据库,涵盖多种技术、方法与系统特征。该数据库不仅为模型提供输入,更使文献本身可被分析。我们揭示了学术假设与实证数据之间的偏差,并展示了研究重点在技术、区域和时间维度上的演变。为促进社区使用,数据库通过交互式仪表板提供,支持用户按需筛选、分析与下载数据。

原文摘要 · Abstract (English)

Energy system models guide societally important decisions, but their credibility rests on quantitative assumptions that are difficult to source and audit. Meta-analyses can improve transparency and modeling practices, but the rapid growth of publications makes manual information extraction increasingly impractical. Consequently, databases are updated infrequently and efforts are often duplicated across research groups. Here, we demonstrate the highly accurate automated extraction of quantitative information from 76,000 energy system studies published since 2010. We compile 3.2 million structured quantitative data points together with 20 million associated metadata entries, spanning a broad spectrum of technologies, methodological approaches and system characteristics. Beyond providing input data for models, the resulting FAIR database make the energy systems literature itself analysable. We show where academic assumptions diverge from empirical observed data, and how research priorities vary at scale across technologies, regions and time. To facilitate broad use within the community, the database is provided through an interactive dashboard, enabling users to filter, analyse and download data according to their specific research needs.

能源建模数据提取自动化数据库

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。