用AI自动更新钙钛矿太阳能电池数据库,让研究跟上论文发布速度。
An autonomous living database for perovskite photovoltaics
- 用大模型+物理验证自动从文献提取数据,精度超90%。
- 发现2021年后领域转向反式结构和甲脒基配方,电压损耗持续下降。
- 适合光伏研究者、数据驱动型创新团队快速追踪前沿。
科学发现因人工整理无法跟上论文爆炸式增长而严重受阻,导致知识鸿沟扩大。在光伏领域尤为突出:主流钙钛矿太阳能电池数据库自2021年起停滞更新,尽管研究产出持续增加。本文提出自主、可自我更新的动态数据库PERLA,通过集成大语言模型与物理感知验证机制,从持续的文献流中提取复杂器件数据,实现人类级精度(>90%),并消除标注者差异。将该系统应用于此前不可及的2021年后文献,揭示出被数据滞后掩盖的关键演化趋势:领域已明确转向采用自组装单层和甲脒基富集成分的反式结构,推动电压损耗持续降低。PERLA使静态论文转化为动态知识资源,使数据驱动发现能以出版速度运行。
原文摘要 · Abstract (English)
Scientific discovery is severely bottlenecked by the inability of manual curation to keep pace with exponential publication rates. This creates a widening knowledge gap. This is especially stark in photovoltaics, where the leading database for perovskite solar cells has been stagnant since 2021 despite massive ongoing research output. Here, we resolve this challenge by establishing an autonomous, self-updating living database (PERLA). Our pipeline integrates large language models with physics-aware validation to extract complex device data from the continuous literature stream, achieving human-level precision (>90%) and eliminating annotator variance. By employing this system on the previously inaccessible post-2021 literature, we uncover critical evolutionary trends hidden by data lag: the field has decisively shifted toward inverted architectures employing self-assembled monolayers and formamidinium-rich compositions, driving a clear trajectory of sustained voltage loss reduction. PERLA transforms static publications into dynamic knowledge resources that enable data-driven discovery to operate at the speed of publication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。