用二进制嵌入实现快速文本检索,让野生动物数据库搜索更高效。
Compact Hypercube Embeddings for Fast Text-based Wildlife Observation Retrieval

- 将文本、图像、音频统一映射到哈明空间,用紧凑二进制编码加速搜索。
- 在iNaturalist2024和iNatSounds2024上表现优于或媲美连续嵌入,内存降低90%以上。
- 适合需要大规模生物多样性数据快速检索的科研与保护系统使用。
大规模生物多样性监测平台依赖多模态野生动物观测数据。尽管近期基础模型能生成丰富的视觉、音频与语言语义表征,但高维相似性搜索带来的计算开销仍使从海量档案中检索相关观测变得困难。本文提出紧凑超立方体嵌入(compact hypercube embeddings),通过紧凑二进制表示实现大规模野生动物图像与音频库的快速文本检索。基于跨视图码对齐哈希框架,将自然语言描述与视觉或声学观测在共享哈明空间中对齐。方法利用预训练的野生动物基础模型(如BioCLIP和BioLingual),采用参数高效微调策略进行哈希适配。在iNaturalist2024(文本到图像)和iNatSounds2024(文本到音频)等大规模基准上评估,结果表明:使用离散超立方体嵌入的检索性能与连续嵌入相当甚至更优,同时大幅降低内存占用与搜索成本。此外,哈希目标持续提升底层编码器表征能力,增强检索效果与零样本泛化性能。这些结果证明,基于二进制的文本检索可为生物多样性监测系统提供可扩展、高效的大型档案搜索能力。
原文摘要 · Abstract (English)
Large-scale biodiversity monitoring platforms increasingly rely on multimodal wildlife observations. While recent foundation models enable rich semantic representations across vision, audio, and language, retrieving relevant observations from massive archives remains challenging due to the computational cost of high-dimensional similarity search. In this work, we introduce compact hypercube embeddings for fast text-based wildlife observation retrieval, a framework that enables efficient text-based search over large-scale wildlife image and audio databases using compact binary representations. Building on the cross-view code alignment hashing framework, we extend lightweight hashing beyond a single-modality setup to align natural language descriptions with visual or acoustic observations in a shared Hamming space. Our approach leverages pretrained wildlife foundation models, including BioCLIP and BioLingual, and adapts them efficiently for hashing using parameter-efficient fine-tuning. We evaluate our method on large-scale benchmarks, including iNaturalist2024 for text-to-image retrieval and iNatSounds2024 for text-to-audio retrieval, as well as multiple soundscape datasets to assess robustness under domain shift. Results show that retrieval using discrete hypercube embeddings achieves competitive, and in several cases superior, performance compared to continuous embeddings, while drastically reducing memory and search cost. Moreover, we observe that the hashing objective consistently improves the underlying encoder representations, leading to stronger retrieval and zero-shot generalization. These results demonstrate that binary, language-based retrieval enables scalable and efficient search over large wildlife archives for biodiversity monitoring systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。