arXiv:2412.02262cs.CVcs.LG2024-12中稿 · NeurIPS被引 1

用检索增强生成技术实现无需训练的海洋生物识别。

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation

  • 基于预训练视觉语言模型与检索增强生成,实现开放域图像分析。
  • 在渔船视频上无需特定训练即实现鱼类分类,准确率显著提升。
  • 适合海洋监测、生态保护等小样本、多变环境的应用场景。

气候变化对海洋生物多样性的破坏正威胁全球依赖健康海洋维生的社区与经济。将计算机视觉应用于海洋保护等特定现实领域时,面临动态多变环境带来的挑战:传统自上而下的学习方法难以应对长尾分布、泛化能力差及领域迁移问题。大规模物种识别尤其困难,因需适应新环境并识别罕见或未见过的物种。为此,我们提出采用自下而上的开放域学习框架,作为海洋应用中图像与视频分析的稳健、可扩展解决方案。初步演示使用预训练视觉语言模型(VLMs)结合检索增强生成(RAG)作为基础,为多种架构、训练与工程优化留出空间。通过在船上渔船视频中分类鱼类的初步应用验证该方法,展示了无需领域特定训练或任务知识即可实现出色的检索与预测能力。

原文摘要 · Abstract (English)

Climate change's destruction of marine biodiversity is threatening communities and economies around the world which rely on healthy oceans for their livelihoods. The challenge of applying computer vision to niche, real-world domains such as ocean conservation lies in the dynamic and diverse environments where traditional top-down learning struggle with long-tailed distributions, generalization, and domain transfer. Scalable species identification for ocean monitoring is particularly difficult due to the need to adapt models to new environments and identify rare or unseen species. To overcome these limitations, we propose leveraging bottom-up, open-domain learning frameworks as a resilient, scalable solution for image and video analysis in marine applications. Our preliminary demonstration uses pretrained vision-language models (VLMs) combined with retrieval-augmented generation (RAG) as grounding, leaving the door open for numerous architectural, training and engineering optimizations. We validate this approach through a preliminary application in classifying fish from video onboard fishing vessels, demonstrating impressive emergent retrieval and prediction capabilities without domain-specific training or knowledge of the task itself.

海洋监测视觉语言模型RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。