arXiv:2503.18548cs.CV2025-03被引 1

评测食物识别中未知类别检测效果,发现虚拟对齐法最优。

Benchmarking Post-Hoc Unknown-Category Detection in Food Recognition

  • 用虚拟对齐法融合输出与特征空间提升检测能力
  • 基于变换器的模型在未知样本检测上显著优于卷积模型
  • 高准确率模型通常具备更强的未知类别识别能力

食物识别模型常难以区分已见与未见样本,易将未见类别误判为已知类别,导致系统级错误。理想情况下,模型应能在遇到未知样本时主动提示用户。由于缺乏真实场景下的研究,本文对多种后处理分布外(OOD)检测方法在细粒度食物识别中的表现进行了实证分析。结果表明,虚拟对齐(ViM)方法整体表现最佳,可能因其结合了输出概率与特征表示。此外,研究再次验证了高精确度模型在各类检测方法中表现更优的规律。变换器架构在多数检测方法中均显著优于卷积模型。

原文摘要 · Abstract (English)

Food recognition models often struggle to distinguish between seen and unseen samples, frequently misclassifying samples from unseen categories by assigning them an in-distribution (ID) label. This misclassification presents significant challenges when deploying these models in real-world applications, particularly within automatic dietary assessment systems, where incorrect labels can lead to cascading errors throughout the system. Ideally, such models should prompt the user when an unknown sample is encountered, allowing for corrective action. Given no prior research exploring food recognition in real-world settings, in this work we conduct an empirical analysis of various post-hoc out-of-distribution (OOD) detection methods for fine-grained food recognition. Our findings indicate that virtual logit matching (ViM) performed the best overall, likely due to its combination of logits and feature-space representations. Additionally, our work reinforces prior notions in the OOD domain, noting that models with higher ID accuracy performed better across the evaluated OOD detection methods. Furthermore, transformer-based architectures consistently outperformed convolution-based models in detecting OOD samples across various methods.

食物识别未知检测视觉模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。