大规模生物图像训练让模型自动生成有用特征,可精准识别物种与生态特性。
BioCLIP 2: Emergent Properties from Scaling Hierarchical Contrastive Learning
- 基于层级对比学习,在2.14亿张生物图像上训练模型
- 模型在物种分类、栖息地预测等任务中表现超预期,准确率显著提升
- 嵌入空间具生物学意义,能区分种间差异与种内变异,适合生物研究者使用
在大规模对比视觉-语言训练下,基础模型展现出超越初始目标的涌现行为。我们通过构建包含2.14亿张生物体图像的TreeOfLife-200M数据集,训练了BioCLIP 2模型以区分不同物种。尽管训练目标狭窄,该模型在物种分类、栖息地判断和性状预测等多类生物视觉任务中表现出极高的准确性。我们发现,模型学习到的嵌入空间中:种间分布紧密对应功能与生态含义(如喙形大小、栖息地);种内变异(如生命周期阶段、性别)不仅未被削弱,反而在与种间区分正交的子空间中得到更好分离。我们通过形式化证明与分析揭示,层级监督与对比目标共同促进这些涌现特性。关键的是,这些特性随训练数据规模增大而愈发显著,形成具有生物学意义的嵌入空间。
原文摘要 · Abstract (English)
Foundation models trained at scale exhibit remarkable emergent behaviors, learning new capabilities beyond their initial training objectives. We find such emergent behaviors in biological vision models via large-scale contrastive vision-language training. To achieve this, we first curate TreeOfLife-200M, comprising 214 million images of living organisms, the largest and most diverse biological organism image dataset to date. We then train BioCLIP 2 on TreeOfLife-200M to distinguish different species. Despite the narrow training objective, BioCLIP 2 yields extraordinary accuracy when applied to various biological visual tasks such as habitat classification and trait prediction. We identify emergent properties in the learned embedding space of BioCLIP 2. At the inter-species level, the embedding distribution of different species aligns closely with functional and ecological meanings (e.g., beak sizes and habitats). At the intra-species level, instead of being diminished, the intra-species variations (e.g., life stages and sexes) are preserved and better separated in subspaces orthogonal to inter-species distinctions. We provide formal proof and analyses to explain why hierarchical supervision and contrastive objectives encourage these emergent properties. Crucially, our results reveal that these properties become increasingly significant with larger-scale training data, leading to a biologically meaningful embedding space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。