在真实植物病害数据上评估了6种异常检测方法,发现基于能量的微调效果最佳。
Beyond Toy Benchmarks: A Systematic Evaluation of OOD Detection Methods For Plant Pathology Classification

- 采用能量模型微调重构特征空间并校准评分函数
- 在多种分布外场景下优于softmax基线,且保持原任务准确率
- 揭示约束优化方法在中等规模数据下的训练不稳定性
分布外(OOD)检测对深度学习系统可靠部署至关重要,但现有方法多在小规模、视觉同质的基准上评估。本文在具有自然分布偏移的细粒度植物病害分类数据集Plant Pathology 2021上,系统评估了六种涵盖后处理评分、辅助目标、能量模型和约束优化的OOD检测方法。结果显示,能量模型微调在各类OOD设置中表现最优,显著提升检测性能的同时保持了原任务准确率。分析表明,性能提升源于嵌入空间重构与评分函数校准的协同作用。此外,我们还发现了约束优化方法在扩展至中等规模数据集时出现的实用训练不稳定性,这一问题在现有文献中鲜有提及。结果表明,在真实世界领域特定数据上实现严谨的OOD检测是可行的,而仅依赖基准评估可能无法捕捉实际应用中的挑战。
原文摘要 · Abstract (English)
Out-of-distribution (OOD) detection is essential for reliable deployment of deep learning systems, yet the majority of existing methods are evaluated on small, visually homogeneous benchmarks. In this work, we study six OOD detection methods spanning post-hoc scoring, auxiliary objectives, energy-based models, and constrained optimization on the Plant Pathology 2021 dataset, a fine-grained task with natural distribution shifts. Energy-based fine-tuning performs best across OOD settings, improving detection over the softmax baseline while preserving in-distribution accuracy. Analysis shows these gains stem from both a restructuring of the embedding space alongside calibration of the scoring function. We further document practical training instabilities that arise when scaling constrained optimization methods to moderate-sized datasets, findings that are largely absent from existing literature. Our results demonstrate that principled OOD detection is achievable on real-world domain-specific data and that benchmark evaluations alone may not capture the challenges that emerge in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。