用合成数据提升零样本真菌分类能力
FungalZSL: Zero-Shot Fungal Classification with Image Captioning Using a Synthetic Data Approach
- 用LLM生成真菌生长阶段描述文本,合成图像增强数据
- 将图文映射到CLIP共享空间,提升跨模态对齐效果
- 适合真菌识别、生物图像分析研究者参考
大视觉语言模型(如CLIP)在零样本分类中的表现依赖于大规模且对齐良好的图文数据集。本文提出两种互补的合成数据源:一是利用LLaMA3.2生成的真菌生长各阶段文本描述,二是多样化的合成真菌图像。通过将这些数据投影至CLIP的共享表示空间,聚焦不同生长阶段的图文对齐,强化零样本分类性能。同时,比较不同LLM生成策略的文本输出,探索知识迁移机制以优化各阶段分类效果。
原文摘要 · Abstract (English)
The effectiveness of zero-shot classification in large vision-language models (VLMs), such as Contrastive Language-Image Pre-training (CLIP), depends on access to extensive, well-aligned text-image datasets. In this work, we introduce two complementary data sources, one generated by large language models (LLMs) to describe the stages of fungal growth and another comprising a diverse set of synthetic fungi images. These datasets are designed to enhance CLIPs zero-shot classification capabilities for fungi-related tasks. To ensure effective alignment between text and image data, we project them into CLIPs shared representation space, focusing on different fungal growth stages. We generate text using LLaMA3.2 to bridge modality gaps and synthetically create fungi images. Furthermore, we investigate knowledge transfer by comparing text outputs from different LLM techniques to refine classification across growth stages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。