评估水体分割模型时,仅用平均指标无法说明性能优劣原因。
A Controlled Evaluation of Model Rankings and Input Reliance in Surface Water Segmentation

- 通过多轮对比测试分析模型配置稳定性
- 发现地形和世界覆盖数据影响预测可靠性
- 提醒评估需匹配结论类型,避免误判
表面水体分割的性能评估通常依赖全局交并比(IoU)等综合指标对模型配置进行排名。然而,这种排名无法说明为何某系统表现更好、排序是否稳定,或预测对单个输入的依赖程度。本文在Sen1Floods11数据集上通过重复配置比较、成对测试芯片分析、固定检查点输入压力测试和地理加权,辅以对GEOID-Flood上监督输入配置的二次评估,发现跨模态学生模型在Sen1Floods11上三种子均IoU最高,但排序随种子和地理权重变化;辅助输入排名在Swin-UNet与U-Net间不一致。GEOID-Flood结果表明监督辅助输入效应具较强一致性,但架构排序仍依赖具体配置。固定检查点测试显示模型依赖地形和WorldCover数据,但未体现明确输入优势;目标语义与后期WorldCover先验限制评估范围为回顾性全水体分割。结果表明,综合指标仍可用于完整配置排名,但排序稳定性、组件归因、输入依赖性和部署范围需独立证据支持。评估应与所声称结论相匹配。
原文摘要 · Abstract (English)
Performance evaluation for surface-water segmentation commonly uses an aggregate metric such as global intersection-over-union (IoU) to rank model configurations. However, a configuration ranking does not by itself establish why one system performs better, whether a close ordering is stable, or how strongly predictions rely on individual inputs. We examine these distinctions primarily on Sen1Floods11 through repeated configuration comparisons, paired test-chip analysis, fixed-checkpoint input stress tests, and geographic reweighting, with a targeted secondary evaluation of supervised input configurations on GEOID-Flood. The cross-modal student achieves the highest three-seed mean IoU on Sen1Floods11, but close orderings vary across seeds and geographic weighting, while ancillary-input rankings differ between Swin-UNet and U-Net. The GEOID-Flood evaluation shows substantial agreement in supervised ancillary-input effects, although the exact architecture ordering remains configuration dependent. Fixed-checkpoint tests further establish reliance on terrain and WorldCover without establishing a clean-input performance benefit, while target semantics and the later WorldCover prior restrict the evaluation to retrospective all-water segmentation. These results show that aggregate metrics remain useful for ranking complete configurations, but ranking stability, component attribution, input reliance, and deployment scope require distinct evidence. Performance evaluation should therefore match the evidence reported to the claim being made.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。