针对科学图像长尾识别难题,提出新框架提升尾部类别准确率。
SciLT: Long-tailed Image Classification under Scientific Image Domains
- 融合中间层与最终层特征,自适应优化表示
- 在三个科学数据集上尾部类别的准确率显著提升
- 适合需迁移基础模型至科学领域的研究者
长尾识别虽受益于基础模型和微调范式,但现有研究和基准多集中于自然图像领域,其预训练与微调数据分布相似。而科学图像具有独特的视觉特征和标注信号,挑战了基础模型微调的有效性。本文在纯视觉微调范式下研究科学领域的长尾识别问题。在三个科学基准上的实验表明,直接微调基础模型收益有限,且中间层特征对尾部类别尤为关键。基于此,提出SciLT框架,通过自适应特征融合与双监督学习,联合利用中间层与最终层特征,在头尾类别间实现平衡性能。大量实验证明,SciLT持续优于现有方法,为科学长尾识别建立了强而实用的基线,并为在存在显著领域偏移的科学数据上适配基础模型提供重要指导。
原文摘要 · Abstract (English)
Long-tailed recognition has benefited from foundation models and fine-tuning paradigms, yet existing studies and benchmarks are mainly confined to natural image domains, where pre-training and fine-tuning data share similar distributions. In contrast, scientific images exhibit distinct visual characteristics and supervision signals, raising questions about the effectiveness of fine-tuning foundation models in such settings. In this work, we investigate scientific long-tailed recognition under a purely visual and fine-tuning paradigm. Experiments on three scientific benchmarks show that fine-tuning foundation models yields limited gains, and reveal that penultimate-layer features play an important role, particularly for tail classes. Motivated by these findings, we propose SciLT, a framework that exploits multi-level representations through adaptive feature fusion and dual-supervision learning. By jointly leveraging penultimate- and final-layer features, SciLT achieves balanced performance across head and tail classes. Extensive experiments demonstrate that SciLT consistently outperforms existing methods, establishing a strong and practical baseline for scientific long-tailed recognition and providing valuable guidance for adapting foundation models to scientific data with substantial domain shifts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。