arXiv:2507.08248cs.CVcs.IR2025-07被引 2

用视觉模型与数据均衡提升真菌细粒度少样本分类准确率

Transfer Learning and Mixup for Fine-Grained Few-Shot Fungi Classification

  • 采用视觉变换器结合平衡采样和数据增强
  • 在FungiTastic数据集上超越基线,排名35/74
  • 适合对少样本细粒度识别感兴趣的研究者

由于真菌物种间细微差异和种内高度变异性,其精准识别在计算机视觉中极具挑战。本文针对FungiCLEF 2025竞赛,聚焦使用FungiTastic Few-Shot数据集的少样本细粒度视觉分类(FGVC)。我们团队(DS@GT)尝试了多种视觉变换器模型、数据增强、加权采样及文本信息融合策略。同时探索生成式AI模型通过结构化提示实现零样本分类,但发现其表现显著低于基于视觉的模型。最终模型在私有测试集上排名35/74(赛后评估),表明元数据选择与领域自适应多模态学习仍有优化空间。代码已开源:https://github.com/dsgt-arc/fungiclef-2025。

原文摘要 · Abstract (English)

Accurate identification of fungi species presents a unique challenge in computer vision due to fine-grained inter-species variation and high intra-species variation. This paper presents our approach for the FungiCLEF 2025 competition, which focuses on few-shot fine-grained visual categorization (FGVC) using the FungiTastic Few-Shot dataset. Our team (DS@GT) experimented with multiple vision transformer models, data augmentation, weighted sampling, and incorporating textual information. We also explored generative AI models for zero-shot classification using structured prompting but found them to significantly underperform relative to vision-based models. Our final model outperformed both competition baselines and highlighted the effectiveness of domain specific pretraining and balanced sampling strategies. Our approach ranked 35/74 on the private test set in post-completion evaluation, this suggests additional work can be done on metadata selection and domain-adapted multi-modal learning. Our code is available at https://github.com/dsgt-arc/fungiclef-2025.

少样本学习细粒度分类视觉变换器真菌识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。