十种模型在高光谱分类中因邻近像素泄露而误判,真实性能被严重夸大。
Ten Architectures, One Error: Shared Failure Modes in Hyperspectral Classification under Spatially Disjoint Evaluation

- 按空间隔离设计评估协议,避免训练测试像素相邻导致的虚假准确率。
- 平均宏F1下降0.147,部分模型排名变动达5位,揭示性能虚高。
- 所有模型错判相似像素,暴露数据本身存在难以克服的光谱歧义。
高光谱图像分类仍严重依赖单场景内随机像素划分。萨利纳斯数据集经随机分割后,是对比不同架构最常用的数据集之一。然而,随机划分下大量测试像素紧邻训练像素,人为抬高报告准确率。本文提出一种无泄漏评估协议,将空间分离与模型感受野相匹配。对十类不同架构(包括经典、光谱、光谱-空间、Transformer、视觉主干和状态空间模型)应用该协议后发现,平均宏F1下降0.147,模型排名最高变动5位。此外,无泄漏评估限制了可测试架构范围——每个划分仅支持有限半径内的补丁,因此报告该半径与感受野至关重要。研究还发现,所有十种架构均错误分类几乎相同的像素,揭示数据中存在未被任何模型解决的光谱歧义。
原文摘要 · Abstract (English)
Hyperspectral image classification still relies heavily on random pixel splits within a single scene. The Salinas dataset, randomly split, is among the most widely used datasets for comparing different architectures. However, under a random split method, a large fraction of test pixels fall immediately adjacent to a training pixel, which inflates reported accuracy. This work introduces a leakage-free evaluation protocol linking spatial separation to the model's receptive field. Applying this protocol across ten different architectures, including classical, spectral, spectral-spatial, transformer, vision-backbone, and state-space families, shows that Macro-F1 drops by 0.147 on average and model rankings change by as many as five places. Furthermore, leakage-free evaluation limits which architectures can be tested on a given benchmark. Since each partition supports patches only within a finite radius, reporting this radius alongside the receptive field is essential for fair comparison. In addition, this study reveals that all ten architectures misclassify largely the same pixels, pointing to a spectral ambiguity in the data that none of them resolves.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。