构建全球最大模糊物种数据集,助力AI识别难以区分的生物种类
CrypticBio: A Large Multimodal Dataset for Visually Confusing Biodiversity
- 聚焦视觉相似物种,整合百万级图像与多模态标注
- 涵盖6.7万种、52000组混淆物种,覆盖濒危与入侵物种
- 适合生态学、计算机视觉与跨学科研究者使用
我们提出CrypticBio,目前最大公开的多模态模糊物种数据集,专为支持生物多样性人工智能研究而设计。该数据集针对视觉上几乎无法区分的物种(即隐匿物种)进行专门构建,涵盖67,000个物种的52,000组独特混淆组合,包含1.66亿张高质量图像。数据包含科学术语、多文化多语言命名、层级分类、时空上下文及关联混淆组等丰富注释,支持多模态人工智能研究。我们还提供了开源数据集构建工具CrypticBio-Curate。数据不仅融合视觉与语言信息,更引入地理与时间维度作为辅助判别线索。我们在常见、未见、濒危和入侵物种子集上测试了多个前沿基础模型,验证了地理背景对视觉-语言零样本学习的关键影响。CrypticBio旨在推动具备真实世界适应能力的生物多样性智能模型发展。
原文摘要 · Abstract (English)
We present CrypticBio, the largest publicly available multimodal dataset of visually confusing species, specifically curated to support the development of AI models in the context of biodiversity applications. Visually confusing or cryptic species are groups of two or more taxa that are nearly indistinguishable based on visual characteristics alone. While much existing work addresses taxonomic identification in a broad sense, datasets that directly address the morphological confusion of cryptic species are small, manually curated, and target only a single taxon. Thus, the challenge of identifying such subtle differences in a wide range of taxa remains unaddressed. Curated from real-world trends in species misidentification among community annotators of iNaturalist, CrypticBio contains 52K unique cryptic groups spanning 67K species, represented in 166 million images. Rich research-grade image annotations--including scientific, multicultural, and multilingual species terminology, hierarchical taxonomy, spatiotemporal context, and associated cryptic groups--address multimodal AI in biodiversity research. For easy dataset curation, we provide an open-source pipeline CrypticBio-Curate. The multimodal nature of the dataset beyond vision-language arises from the integration of geographical and temporal data as complementary cues to identifying cryptic species. To highlight the importance of the dataset, we benchmark a suite of state-of-the-art foundation models across CrypticBio subsets of common, unseen, endangered, and invasive species, and demonstrate the substantial impact of geographical context on vision-language zero-shot learning for cryptic species. By introducing CrypticBio, we aim to catalyze progress toward real-world-ready biodiversity AI models capable of handling the nuanced challenges of species ambiguity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。