arXiv:2510.20095cs.CVcs.CL2025-10被引 3

用合成描述句提升生物模型对图像的理解能力

BioCAP: Exploiting Synthetic Captions Beyond Labels in Biological Foundation Models

  • 用多模态大模型生成精准的物种描述性文本
  • 训练出在分类和图文检索上表现优异的BioCAP模型
  • 适合做生物图像理解与跨模态检索的研究者

本研究探索将描述性文本作为生物多模态基础模型的额外监督信号。图像与文本可视为物种潜在形态空间中的互补样本,各自捕捉特定生物特征。训练中引入文本有助于对齐共享潜在结构,突出可能具有诊断意义的特征,同时抑制虚假关联。主要挑战在于大规模获取准确、实例相关的文本描述。为此,本文利用维基百科提供的视觉信息和分类定制的格式示例,通过多模态大语言模型生成合成描述文本,有效降低幻觉,获得真实、实例化的描述。基于这些文本,训练了BioCAP(即BioCLIP with Captions)模型,在物种分类与文本-图像检索任务中均取得优异性能,验证了描述性文本在连接生物图像与多模态基础模型中的价值。

原文摘要 · Abstract (English)

This work investigates descriptive captions as an additional source of supervision for biological multimodal foundation models. Images and captions can be viewed as complementary samples from the latent morphospace of a species, each capturing certain biological traits. Incorporating captions during training encourages alignment with this shared latent structure, emphasizing potentially diagnostic characters while suppressing spurious correlations. The main challenge, however, lies in obtaining faithful, instance-specific captions at scale. This requirement has limited the utilization of natural language supervision in organismal biology compared with many other scientific domains. We complement this gap by generating synthetic captions with multimodal large language models (MLLMs), guided by Wikipedia-derived visual information and taxon-tailored format examples. These domain-specific contexts help reduce hallucination and yield accurate, instance-based descriptive captions. Using these captions, we train BioCAP (i.e., BioCLIP with Captions), a biological foundation model that captures rich semantics and achieves strong performance in species classification and text-image retrieval. These results demonstrate the value of descriptive captions beyond labels in bridging biological images with multimodal foundation models.

生物图像多模态合成数据基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。