用碳排放量化机器学习加速光伏材料发现的环保效益
The carbon cost of materials discovery: Can machine learning really accelerate the discovery of new photovoltaics?
- 用机器学习替代部分密度泛函计算,降低碳排放
- 直接预测效率比中间生成光谱更省资源且准确
- 为绿色高通量材料筛选提供可量化的碳成本框架
计算筛选已成为发现高性能光伏材料的重要手段,主流流程依赖密度泛函理论(DFT)估算与太阳能转换相关的电子和光学性质。尽管相比实验方法更高效,但DFT仍带来显著的计算与环境成本。近年来,机器学习(ML)模型作为DFT替代方案受到关注,能在保持预测性能的同时大幅减少资源消耗。本研究复现了典型的基于DFT的工作流,用于估算材料最大效率极限,并逐步用ML代理模型替换其组件。通过量化各策略的二氧化碳排放,评估了预测精度与环境成本之间的权衡。结果表明,多种混合型ML/DFT策略可在精度-排放前沿实现优化。直接预测效率等标量值比以预测吸收光谱为中间步骤更为可行。有趣的是,基于DFT数据训练的ML模型在筛选任务中甚至优于使用其他交换-关联泛函的DFT流程,凸显数据驱动方法的一致性与实用性。我们还评估了扩展数据集和针对光伏特征优化模型架构对提升筛选效果的作用。本工作为构建低碳、高通量材料发现管道提供了定量框架。
原文摘要 · Abstract (English)
Computational screening has become a powerful complement to experimental efforts in the discovery of high-performance photovoltaic (PV) materials. Most workflows rely on density functional theory (DFT) to estimate electronic and optical properties relevant to solar energy conversion. Although more efficient than laboratory-based methods, DFT calculations still entail substantial computational and environmental costs. Machine learning (ML) models have recently gained attention as surrogates for DFT, offering drastic reductions in resource use with competitive predictive performance. In this study, we reproduce a canonical DFT-based workflow to estimate the maximum efficiency limit and progressively replace its components with ML surrogates. By quantifying the CO$_2$ emissions associated with each computational strategy, we evaluate the trade-offs between predictive efficacy and environmental cost. Our results reveal multiple hybrid ML/DFT strategies that optimize different points along the accuracy--emissions front. We find that direct prediction of scalar quantities, such as maximum efficiency, is significantly more tractable than using predicted absorption spectra as an intermediate step. Interestingly, ML models trained on DFT data can outperform DFT workflows using alternative exchange--correlation functionals in screening applications, highlighting the consistency and utility of data-driven approaches. We also assess strategies to improve ML-driven screening through expanded datasets and improved model architectures tailored to PV-relevant features. This work provides a quantitative framework for building low-emission, high-throughput discovery pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。