arXiv:2609.03985cs.CVcs.CL2026-09

零样本模型识鱼准率达72%,但命名和语言影响极大。

IchthyoNoma: Nomenclature and Context Sensitivity of Zero-Shot Biological Vision--Language Models for Bangladeshi Freshwater Fish Recognition

论文配图:IchthyoNoma: Nomenclature and Context Sensitivity of Zero-Shot Biological Vision--Language Models for Bangladeshi Freshwater Fish Recognition
图 1 · 摘自论文原文
  • 用生物专化与多语言提示提升鱼类识别性能
  • 英文名下最高准确率72.36%,孟加拉语提示接近随机
  • 适合关注跨语言生物识别的科研人员

零样本视觉-语言模型(VLMs)被广泛用于无训练物种识别,但其准确率不仅反映视觉知识。我们在两个孟加拉淡水鱼数据集(共10,321张图像)上评估CLIP、BioCLIP、BioCLIP2及多语言Jina CLIP v2。BioCLIP2在BFF-15上使用英文常见名达72.36%准确率,在SylFishBD上用学名达68.91%,而通用CLIP仅25.15%和14.40%。使用孟加拉语提示时,平衡准确率仅14.22%-14.29%;Jina部分恢复至21.89%和16.36%,但仅用孟加拉语名称仍为14.29%。干预实验显示:弱模糊无显著影响,强模糊/灰度遮蔽导致小幅下降,白化遮蔽影响较大,且结果高度依赖物种。零样本生物VLM表现受生物学专化性、多语言对齐、命名体系、提示设计和上下文共同影响。

原文摘要 · Abstract (English)

Zero-shot vision-language models (VLMs) are increasingly used as training-free species recognizers, but reported accuracy can reflect more than visual species knowledge. We audit CLIP, BioCLIP, BioCLIP2, and a multilingual Jina CLIP v2 control on seven freshwater-fish categories from two Bangladeshi sources (10,321 images). BioCLIP2 reaches 72.36% on BFF-15 with English common names and 68.91% on SylFishBD with scientific names, versus 25.15% and 14.40% for generic CLIP. BioCLIP2 Bengali prompts are near chance in balanced accuracy (14.22-14.29%); Jina partially recovers Bengali discrimination to 21.89% and 16.36%, but bare Bengali names return to 14.29% on both sources. Paired SylFishBD interventions show no significant weak-blur effect, modest losses from stronger blur/gray masking, a larger white-mask artifact, and strong species dependence. Zero-shot biological VLM scores therefore jointly reflect biological specialization, multilingual alignment, nomenclature, prompt formulation, and context.

零样本识别生物视觉多语言模型鱼类分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。