用可解释模型揭示野火风险的本地驱动因素及跨县差异
WildfireGenome: Interpretable Machine Learning Reveals Local Drivers of Wildfire Risk and Their Cross-County Variation
- 融合7类联邦野火指标,构建高分辨率综合风险标签
- 模型准确率0.755-0.878,主成分解释87%-94%方差
- 揭示针叶林占比30%-40%时风险剧增,适合规划决策者
当前野火风险评估依赖粗粒度灾害图与不可解释的机器学习模型,虽提升区域准确性却牺牲了决策尺度的可解释性。WildfireGenome通过三部分解决:(1) 将七类联邦野火指标融合为基于PCA的符号对齐复合风险标签,空间分辨率达H3 Level-8;(2) 使用随机森林进行局部野火风险分类;(3) 利用SHAP与ICE/PDP分析揭示各县非线性驱动关系。在七个生态多样性的美国县中,模型准确率0.755-0.878,加权二次卡帕系数达0.951,主成分解释87%-94%指标方差。迁移测试显示,在生态相似区域表现稳定,但在差异显著区域失效。解释结果一致指出针叶林覆盖和高程为关键驱动因素,风险在针叶林占比30%-40%时急剧上升。WildfireGenome将野火风险评估从区域预测推进至可解释、决策级分析,助力植被管理、分区规划与基础设施布局。
原文摘要 · Abstract (English)
Current wildfire risk assessments rely on coarse hazard maps and opaque machine learning models that optimize regional accuracy while sacrificing interpretability at the decision scale. WildfireGenome addresses these gaps through three components: (1) fusion of seven federal wildfire indicators into a sign-aligned, PCA-based composite risk label at H3 Level-8 resolution; (2) Random Forest classification of local wildfire risk; and (3) SHAP and ICE/PDP analyses to expose county-specific nonlinear driver relationships. Across seven ecologically diverse U.S. counties, models achieve accuracies of 0.755-0.878 and Quadratic Weighted Kappa up to 0.951, with principal components explaining 87-94% of indicator variance. Transfer tests show reliable performance between ecologically similar regions but collapse across dissimilar contexts. Explanations consistently highlight needleleaf forest cover and elevation as dominant drivers, with risk rising sharply at 30-40% needleleaf coverage. WildfireGenome advances wildfire risk assessment from regional prediction to interpretable, decision-scale analytics that guide vegetation management, zoning, and infrastructure planning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。