首个珊瑚礁图像问答数据集,助力智能保护海洋生态
CoralVQA: A Large-Scale Visual Question Answering Dataset for Coral Reef Image Understanding
- 联合生物学家构建半自动标注流程,确保数据专业性
- 含1.28万张珊瑚图像与27.7万组问答对,覆盖67个珊瑚属
- 为珊瑚健康评估提供新基准,适合生态与AI交叉研究者
珊瑚礁是重要但脆弱的生态系统,需持续监测以支持保护工作。尽管珊瑚图像蕴含关键信息,但其解读依赖领域专业知识。视觉问答(VQA)结合大视觉语言模型(LVLMs),有望实现用户友好的图像交互。然而,将VQA应用于珊瑚图像需解决两个核心挑战:领域特定标注和多维度问题。本文提出CoralVQA,首个面向珊瑚礁分析的大规模视觉问答数据集。包含来自3个大洋的67个珊瑚属共12,805张真实珊瑚图像,以及277,653组问答对,全面评估生态与健康状况。通过与海洋生物学家合作,开发半自动数据构建流程,在保证可扩展性的同时确保专业级数据质量。CoralVQA带来全新挑战,为珊瑚图像中视觉-语言推理研究提供全面基准。对多个前沿LVLM的评估揭示了关键局限与机遇,为未来模型发展奠定基础,尤其聚焦于支持珊瑚保护行动。
原文摘要 · Abstract (English)
Coral reefs are vital yet vulnerable ecosystems that require continuous monitoring to support conservation. While coral reef images provide essential information in coral monitoring, interpreting such images remains challenging due to the need for domain expertise. Visual Question Answering (VQA), powered by Large Vision-Language Models (LVLMs), has great potential in user-friendly interaction with coral reef images. However, applying VQA to coral imagery demands a dedicated dataset that addresses two key challenges: domain-specific annotations and multidimensional questions. In this work, we introduce CoralVQA, the first large-scale VQA dataset for coral reef analysis. It contains 12,805 real-world coral images from 67 coral genera collected from 3 oceans, along with 277,653 question-answer pairs that comprehensively assess ecological and health-related conditions. To construct this dataset, we develop a semi-automatic data construction pipeline in collaboration with marine biologists to ensure both scalability and professional-grade data quality. CoralVQA presents novel challenges and provides a comprehensive benchmark for studying vision-language reasoning in the context of coral reef images. By evaluating several state-of-the-art LVLMs, we reveal key limitations and opportunities. These insights form a foundation for future LVLM development, with a particular emphasis on supporting coral conservation efforts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。