标注粒度影响病害检测模型的捷径学习,但不是根本原因。
Shortcut Learning in a Public Grape Disease Dataset: Annotation Granularity as a Modulator, Not a Cause
- 发现标注粒度差异导致模型对特定类别的错误依赖
- 细粒度标注类别的跨物种误检率高达65.7%,是其标注占比的13.41倍
- 调整标注粒度可削弱或放大捷径,但无法创造捷径
农业病害检测公开数据集常以指标表现判断可用性,却忽略标注方案的一致性。在包含3288张图像、11995个边界框、6个类别的公开葡萄病害数据集上,不同模型容量、输入分辨率和检测范式下,测试集mAP50波动范围与种子间噪声相当,瓶颈集中在小目标上。问题源于数据:一类标注为整叶级别(中位框面积占图像43.16%),其余五类为病斑级别。在5156张非葡萄图像上,65.7%的误检框落入该类别,比其训练标注占比高13.41倍。反事实重训练显示,仅缩小该类标注框,其跨物种误检减少66%;安慰剂对照证实效应特异于该类。相反操作(将最细粒度类粗化至整叶级,框数与标注占比匹配,且分布内准确率更高)未引发误检,而原类仍占50%误检。因此,标注粒度是捷径的调节因子而非成因:可放大或抑制已有缺陷,但无法生成。我们提出无需图像或训练的粒度筛选统计量,并证明空中病斑级检测光学上不可达。该失败模式对分布内评估不可见。
原文摘要 · Abstract (English)
Public datasets for agricultural disease detection are usually judged fit for use from reported metrics, which say nothing about whether the annotation scheme is internally consistent. On one public grape disease dataset (3288 images, 11995 boxes, 6 classes), varying model capacity, input resolution and detection paradigm yields a test-set mAP50 range comparable to seed-to-seed noise, with the bottleneck at small objects across all five architectures. The finding lies on the data side: one class is annotated at whole-leaf level (median box area 43.16% of the image) while the other five are annotated at lesion level. On 5156 cross-species images containing no grape, 65.7% of the false-positive boxes fall into that one class, an over-representation of 13.41x relative to its share of the training annotations. Counterfactual retraining establishes a causal effect of granularity on the magnitude of the shortcut: shrinking only that class's boxes cuts its cross-species false positives by 66%, and a placebo control confirms the effect is specific to the manipulated class. A manipulation in the opposite direction, with criteria registered in advance, returns a negative result: coarsening the finest class to whole-leaf level (0.57% to 40.37%), matched in box count and share of annotations and with higher in-distribution AP, still leaves its cross-species false positives at zero boxes, while the unmanipulated original class holds 50.0% of them. Annotation granularity is therefore a modulator of this shortcut, not its cause: it can amplify or attenuate a sink that already exists, but cannot create one, and what fixes the destination remains open. We also give a granularity screening statistic requiring neither images nor training, and show airborne lesion-level detection to be optically out of reach. The failure mode is invisible to in-distribution evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。