构建大规模海洋多模态数据集,推动AI在海洋科学中的应用。
OceanPile: A Large-Scale Multimodal Ocean Corpus for Foundation Models

- 整合声呐、水下图像等多源数据,建立统一海洋语料库。
- 基于知识图谱生成高质量指令数据,提升模型理解能力。
- 提供人工标注评测基准,适合海洋AI研究者使用。
广袤而未充分探索的海洋在全球气候调节和海洋生物多样性支持中发挥关键作用,但人工智能在此领域影响有限,根源在于数据瓶颈。海洋数据分散于多个来源,具有多模态、高噪声、弱标注等特点,缺乏统一模式与语义对齐。尽管多模态大语言模型(MLLMs)在通用领域取得显著进展,但其在海洋科学中的应用仍受限于缺乏针对海洋环境的大规模、良好对齐的多模态数据集。为此,我们提出OceanPile,一个面向海洋基础模型的大型多模态语料库。包含三个核心部分:OceanCorpus,整合来自权威来源的声呐数据、水下影像、海洋科学视觉内容及科学文本的统一语料;OceanInstruction,通过基于层级海洋概念知识图谱的新颖管道生成的高质量指令数据集;OceanBenchmark,用于严格评估的精标人工评测基准。我们建立了多阶段质量控制流程,确保跨模态科学有效性和一致性。实验验证表明,在该数据上训练的模型性能显著提升。所有数据集均已公开,以推动海洋人工智能领域发展,赋能特定领域MLLMs。
原文摘要 · Abstract (English)
The vast and underexplored ocean plays a critical role in regulating global climate and supporting marine biodiversity, yet artificial intelligence has so far delivered limited impact in this domain due to a fundamental data bottleneck. Specifically, ocean data are highly fragmented across disparate sources and inherently exhibit multi-modal, high-noise, and weakly labeled characteristics, lacking unified schemas and semantic alignment. Although Multimodal Large Language Models (MLLMs) have achieved remarkable success in general domains, their application to ocean science remains severely constrained by the absence of large-scale, well-aligned multimodal datasets tailored to marine environments. To bridge this gap, we introduce OceanPile, a large-scale multimodal corpus designed for ocean foundation models. It comprises three key components: OceanCorpus, a unified collection integrating sonar data, underwater imagery, marine science visuals, and scientific text from diverse authoritative sources; OceanInstruction, a high-quality instruction dataset synthesized via a novel pipeline guided by a hierarchical Ocean Concept Knowledge Graph; and OceanBenchmark, a manually curated evaluation benchmark for rigorous assessment. We establish a multi-stage quality control process to ensure scientific validity and alignment across modalities. Experimental validation demonstrates significant performance improvements for models trained on our data. All datasets are publicly released to advance the field of marine artificial intelligence and empower domain-specific MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。