构建了迄今最大的浮游生物图像数据集,助力海洋生态研究。
Planktonzilla: Multimodal dataset and models for understanding plankton ecosystems

- 整合13种成像系统数据,建立1740万张标准化图像库。
- 用分类层级作为文本训练,监督学习性能优于CLIP式方法。
- 揭示现有生物基础模型在海洋图像中的泛化局限,适合生态与计算机交叉研究者。
海洋浮游生物支撑水生食物网并参与全球碳封存,可靠物种识别对理解海洋健康与气候反馈至关重要。现有分类模型在单一数据集上表现良好,但在不同仪器和环境下泛化能力差,源于孤立的训练数据与标签不一致。为此,我们推出Planktonzilla-17M,一个统一数据集,整合公开的浮游生物图像,涵盖13种成像系统。该数据集包含1740万张图像,标准化分类与地理环境元数据,其中374万张为浮游生物图像,覆盖602个分类阶元(201个达物种级),是目前最大最全面的浮游生物图像数据集。基于此,我们在共享ViT骨干网络上对比监督学习与CLIP风格图文训练。结果表明,以分类谱系为文本时,监督分类器性能可媲美或超越CLIP式训练。进一步发现,BioCLIP与BioCLIP2在零样本和少样本设置下表现不佳。利用Planktonzilla-17M显著提升分类性能,凸显当前生物基础模型在海洋成像领域的局限性。
原文摘要 · Abstract (English)
Marine plankton underpin aquatic food webs and play a key role in global CO2 sequestration, making reliable species identification critical for understanding ocean health and climate feedbacks. Existing classification models perform well on individual collections but fail to generalize across instruments and environments due to isolated training datasets and inconsistent labels. To address this, we introduce Planktonzilla-17M, a unified dataset consolidating publicly available plankton image collections spanning thirteen imaging systems. It comprises 17.4 million images with standardized taxonomy and geo-environmental metadata, including 3.74 million plankton images spanning over 602 taxonomic classes, of which 201 are identified at the species level, making it the largest and most comprehensive plankton image dataset to date. Using this large-scale dataset, we perform a controlled comparison between supervised and CLIP-style image--text training on a shared ViT backbone. We find that a supervised classifier matches or exceeds CLIP-style training when trained using taxonomic lineage as text. We further observe that BioCLIP and BioCLIP2 perform poorly on plankton in zero-shot and few-shot settings. Leveraging Planktonzilla-17M improves plankton classification performance, highlighting the limitations of current biological foundation models in marine imaging domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。