构建海上精细感知与复杂推理的评测基准,填补真实海洋场景研究空白。
MARINER: A 3E-Driven Benchmark for Fine-Grained Perception and Complex Reasoning in Open-Water Environments
- 基于3E范式构建多源海事图像数据集
- 涵盖63类船只、5类动态事故,支持细粒度识别与问答
- 揭示大模型在复杂海况下仍难完成因果推理
由于缺乏专用评测基准,真实开放水域中的细粒度视觉理解与高层推理仍鲜有研究。我们提出MARINER,一个基于新型实体-环境-事件(3E)范式的综合性评测基准。该基准包含16,629张多源海事图像,涵盖63种细粒度船舶类别、多样恶劣环境及5类典型动态海上事件,覆盖细粒度分类、目标检测与视觉问答任务。我们在主流多模态大语言模型(MLLMs)上开展广泛评估并建立基线,发现即使先进模型在复杂海景中仍难以实现细粒度区分与因果推理。作为专用海事评测基准,MARINER填补了真实海洋场景下认知级评估的空白,推动鲁棒视觉-语言模型在开放水域应用中的发展。附录与补充材料见https://lxixim.github.io/MARINER。
原文摘要 · Abstract (English)
Fine-grained visual understanding and high-level reasoning in real-world open-water environments remain under-explored due to the lack of dedicated benchmarks. We introduce MARINER, a comprehensive benchmark built under the novel Entity-Environment-Event (3E) paradigm. MARINER contains 16,629 multi-source maritime images with 63 fine-grained vessel categories, diverse adverse environments, and 5 typical dynamic maritime incidents, covering fine-grained classification, object detection, and visual question answering tasks. We conduct extensive evaluations on mainstream Multimodal Large language models (MLLMs) and establish baselines, revealing that even advanced models struggle with fine-grained discrimination and causal reasoning in complex marine scenes. As a dedicated maritime benchmark, MARINER fills the gap of realistic and cognitive-level evaluation for maritime multimodal understanding, and promotes future research on robust vision-language models for open-water applications. Appendix and supplementary materials are available at https://lxixim.github.io/MARINER.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。