arXiv:2411.00683cs.CVcs.AI2024-11中稿 · WACV 2025被引 47

构建跨模态物种嵌入空间,助力生态研究

TaxaBind: A Unified Embedding Space for Ecological Applications

  • 以物种图像为锚点,融合六类数据构建统一嵌入空间
  • 在8000个样本的TaxaBench-8k上实现零样本分类与跨模态检索
  • 适合生态学家、计算机视觉研究者使用

我们提出TaxaBind,一个用于表征任意物种的统一嵌入空间。该空间融合六种模态:物种的地面图像、地理位置、卫星图像、文本、音频和环境特征,可解决多种生态问题。为学习联合嵌入空间,我们以物种的地面图像作为绑定模态,并提出多模态分块技术,有效将不同模态知识提炼到绑定模态中。我们构建了两个大规模预训练数据集:iSatNat(包含物种图像与卫星图像)和iSoundNat(包含物种图像与音频)。此外,我们引入TaxaBench-8k,一个包含六类配对模态的多样化多模态数据集,用于评估深度学习模型在生态任务上的表现。实验表明,TaxaBind在物种分类、跨模态检索和音频分类等任务上具备强大的零样本与涌现能力。相关数据集与模型已公开于https://github.com/mvrl/TaxaBind。

原文摘要 · Abstract (English)

We present TaxaBind, a unified embedding space for characterizing any species of interest. TaxaBind is a multimodal embedding space across six modalities: ground-level images of species, geographic location, satellite image, text, audio, and environmental features, useful for solving ecological problems. To learn this joint embedding space, we leverage ground-level images of species as a binding modality. We propose multimodal patching, a technique for effectively distilling the knowledge from various modalities into the binding modality. We construct two large datasets for pretraining: iSatNat with species images and satellite images, and iSoundNat with species images and audio. Additionally, we introduce TaxaBench-8k, a diverse multimodal dataset with six paired modalities for evaluating deep learning models on ecological tasks. Experiments with TaxaBind demonstrate its strong zero-shot and emergent capabilities on a range of tasks including species classification, cross-model retrieval, and audio classification. The datasets and models are made available at https://github.com/mvrl/TaxaBind.

生态计算多模态嵌入空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。