评估野火风险不能只看预测准确率,要考察风险信号与实际应对压力是否同步上升。
Risk Is Not the Target: A Monotonic Framework for Evaluating Wildfire Operational Risk Signals

- 提出单调性评估框架,检验风险评分升高是否对应真实救援压力增加。
- 实测显示专家系统在单调性上最优,而深度学习模型虽局部精准但分布不均。
- 强调风险模型价值在于解释实际操作动态,而非单纯预测火灾发生。
使用标准机器学习指标(如F1分数或IoU)评估野火风险系统存在根本缺陷:这些指标衡量的是事件预测准确性,而非连续风险信号的操作一致性。本文提出一种新颖的单调性评估框架,用于检验预测风险评分的升高是否始终对应观测到的操作负荷增加,例如火灾数量、干预时间和部署资源。在法国滨海阿尔卑斯省的实验中,对比了三种结构不同的方法:基于专家的DFE指数、基于GRU的预测模型,以及结合预测AI与大语言模型推理的混合多智能体系统FARS。结果表明,尽管DFE在分类指标上表现不佳,但在整个风险尺度上展现出最均衡的单调性;GRU模型具备较强的局部单调性,但无法生成分布合理的风险等级;FARS继承并暴露了上游信号的结构性局限,未能加以纠正。核心发现是范式转变:优秀的风险模型并非精确预测火灾,而是其序数尺度能有意义地解释实际操作动态,本研究已证实这一点。单调性评估框架代码已开源。
原文摘要 · Abstract (English)
Evaluating wildfire risk systems using standard machine-learning metrics such as F1-score or IoU is fundamentally flawed: these metrics assess event prediction accuracy, not the operational coherence of a continuous risk signal. This work proposes a novel monotonic evaluation framework that measures whether increases in a predicted risk score consistently correspond to increases in observed operational load, such as number of fires, intervention time, and deployed resources. Moreover, we compare three structurally different approaches on the French Alpes-Maritimes department: the expert-based DFE index, GRU- based predictive models, and FARS, a hybrid multi-agent system combining predictive AI with LLM-based reasoning. Experimental results reveal that the DFE, despite poor classification metrics, exhibits the most balanced monotonic behavior across the full risk scale. GRU models achieve strong local monotonicity but fail to produce well-distributed risk levels. FARS inherits and reveals the structural limitations of upstream signals rather than correcting them. The central finding is a paradigm shift: a good risk model does not predict fires accurately, but one whose ordinal scale meaningfully explains operational dynamics, as proved in this paper. Code of the monotonic framework is available on github.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。