arXiv:2603.23883cs.CV2026-03

构建跨模态生物数据集与模型,实现图像、文本、声音的统一表征。

BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignment

  • 基于14,133种生物构建含230万图像和130万音频的大规模数据集。
  • 两阶段训练使音频与图文在物种级语义上对齐,支持跨模态检索。
  • 覆盖三种分类层级的双向检索,助力生态多样性研究。

从多模态数据理解动物物种是计算机视觉与生态学交叉领域的新兴挑战。尽管近期生物模型如BioCLIP已在图像与文本分类信息间实现强对齐,但音频模态的整合仍待解决。本文提出BioVITA,一个面向生物应用的视觉-文本-声音对齐框架,包含:(i) 大规模训练数据集,(ii) 表示模型,(iii) 跨模态检索基准。首先,构建涵盖14,133个物种、含230万张图像与130万段音频的标注数据集,每个物种附带34个生态性状标签。其次,基于BioCLIP2,设计两阶段训练流程,有效对齐音频与视觉、文本表示。第三,开发覆盖三模态所有方向(如图像→音频、音频→文本等)的跨模态检索基准,支持家族、属、种三级分类层级。大量实验表明,该模型学习到的统一表征空间不仅捕捉分类学语义,更蕴含物种级深层语义,推动多模态生物多样性认知。项目主页:https://dahlian00.github.io/BioVITA_Page/

原文摘要 · Abstract (English)

Understanding animal species from multimodal data poses an emerging challenge at the intersection of computer vision and ecology. While recent biological models, such as BioCLIP, have demonstrated strong alignment between images and textual taxonomic information for species identification, the integration of the audio modality remains an open problem. We propose BioVITA, a novel visual-textual-acoustic alignment framework for biological applications. BioVITA involves (i) a training dataset, (ii) a representation model, and (iii) a retrieval benchmark. First, we construct a large-scale training dataset comprising 1.3 million audio clips and 2.3 million images, covering 14,133 species annotated with 34 ecological trait labels. Second, building upon BioCLIP2, we introduce a two-stage training framework to effectively align audio representations with visual and textual representations. Third, we develop a cross-modal retrieval benchmark that covers all possible directional retrieval across the three modalities (i.e., image-to-audio, audio-to-text, text-to-image, and their reverse directions), with three taxonomic levels: Family, Genus, and Species. Extensive experiments demonstrate that our model learns a unified representation space that captures species-level semantics beyond taxonomy, advancing multimodal biodiversity understanding. The project page is available at: https://dahlian00.github.io/BioVITA_Page/

多模态生物识别音频对齐生态数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。