构建3700万条数据的加州火灾预测数据库,助力AI提前预警重大火灾。
California Wildfire Inventory (CAWFI): An Extensive Dataset for Predictive Techniques based on Artificial Intelligence
- 整合2012-2018年加州火灾数据与多类环境指标
- 用2012-2017年数据训练模型,可预测85.7%超30万英亩火灾
- 适合灾害预警、气候建模及地理信息研究者使用
受气候变化和全球生态系统破坏影响,野火正日益威胁环境、基础设施与人类生命。若不采取预防措施,损失将持续扩大。尽管人工智能在野火管理中取得进展,但现有方案多聚焦于火灾发生后的检测。高精度预测需依赖大规模数据集训练机器学习模型。本文提出加州野火清单(CAWFI),包含超过3700万条数据点,用于构建与训练野火预测模型,从而在火灾爆发前进行干预,防止巨灾与突燃火灾。该数据集整合了2012至2018年的每日历史野火数据,以及2012至2022年的各类指标数据:包括气象条件等先行指标、环境变化等滞后指标,以及植被与高程等地质指标,用于评估火灾风险与蔓延模式。使用2012-2017年指标数据训练的时空人工智能模型,成功预测了85.7%的未来火灾,其规模超过30万英亩。该数据集旨在推动野火预测研究,并为其他地区建立类似数据库树立范例。
原文摘要 · Abstract (English)
Due to climate change and the disruption of ecosystems worldwide, wildfires are increasingly impacting environment, infrastructure, and human lives globally. Additionally, an exacerbating climate crisis means that these losses would continue to grow if preventative measures are not implemented. Though recent advancements in artificial intelligence enable wildfire management techniques, most deployed solutions focus on detecting wildfires after ignition. The development of predictive techniques with high accuracy requires extensive datasets to train machine learning models. This paper presents the California Wildfire Inventory (CAWFI), a wildfire database of over 37 million data points for building and training wildfire prediction solutions, thereby potentially preventing megafires and flash fires by addressing them before they spark. The dataset compiles daily historical California wildfire data from 2012 to 2018 and indicator data from 2012 to 2022. The indicator data consists of leading indicators (meteorological data correlating to wildfire-prone conditions), trailing indicators (environmental data correlating to prior and early wildfire activity), and geological indicators (vegetation and elevation data dictating wildfire risk and spread patterns). CAWFI has already demonstrated success when used to train a spatio-temporal artificial intelligence model, predicting 85.7% of future wildfires larger than 300,000 acres when trained on 2012-2017 indicator data. This dataset is intended to enable wildfire prediction research and solutions as well as set a precedent for future wildfire databases in other regions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。