arXiv:2605.04323cs.LGcs.DB2026-05

构建超大规模土壤多模态数据集,助力环境科学智能建模

LUCAS-MEGA: A Large-Scale Multimodal Dataset for Representation Learning in Soil-Environment Systems

论文配图:LUCAS-MEGA: A Large-Scale Multimodal Dataset for Representation Learning in Soil-Environment Systems
图 1 · 摘自论文原文
  • 通过多智能体融合系统整合68个数据源,构建7万样本多模态数据集
  • 预训练模型在特征掩码下表现稳定,可还原已知土壤过程关系
  • 适合土壤科学、环境建模与机器学习交叉研究者使用

理解土壤对农业、碳循环和环境可持续性至关重要,但现有研究受限于分散异构的数据集,难以开展高维表征学习。本文提出LUCAS-MEGA,一个基于欧洲土壤环境观测的大型多模态数据集,以LUCAS调查为核心。该数据集包含超过7万条样本、1000多个特征,涵盖物理、化学、环境、生物及视觉属性,来自68个数据源。为实现规模化集成,我们开发了SoilFuser——一种人机协同的多智能体数据融合管道,统一格式与测量协议,修正单位不一致、编码错误等异常值,融入自然语言注释,并将多模态属性与元数据对齐至统一机器学习可用空间。数据集保留真实土壤观测的关键特性:多模态性、特征覆盖不均与异质不确定性。为验证实用性,我们使用自监督特征掩码训练了多模态表格变压器SoilFormer,实现稳定训练与强预测性能,所学表征支持不确定性感知预测,并能恢复已知土壤过程关系。LUCAS-MEGA已开源,配套可组合、代理友好的API,支持结构化查询与数据驱动工作流。

原文摘要 · Abstract (English)

Understanding soil is fundamental to agriculture, carbon cycling, and environmental sustainability, yet progress is limited by fragmented and heterogeneous datasets that constrain modeling to small-scale predictive settings rather than high-dimensional representation learning. We introduce LUCAS-MEGA, a large-scale multimodal dataset constructed through systematic data fusion of European soil-environment observations, with the LUCAS survey as its backbone. The fused dataset comprises over 70,000 samples and more than 1,000 features spanning physical, chemical, environmental, biological, and visual attributes, aggregated from 68 source datasets. To enable integration at scale, we develop SoilFuser, a multi-agent, human-in-the-loop data fusion pipeline that standardizes heterogeneous data formats and measurement protocols, resolves inconsistencies and invalid entries (e.g., unit inconsistencies, codebook mismatches, and erroneous values), incorporates natural language annotations, and harmonizes multimodal attributes and metadata into a unified, machine learning-ready feature space. The resulting dataset captures key characteristics of real-world soil observations, including multimodality, uneven feature coverage, and heterogeneous uncertainty. To demonstrate the usability of LUCAS-MEGA for data-driven modeling, we pretrain a multimodal tabular transformer (SoilFormer) using a self-supervised objective based on feature masking, achieving stable training, strong predictive performance, and representations that support uncertainty-aware prediction. We further show that the learned representations recover relationships consistent with established soil processes. LUCAS-MEGA is released with open access and is accompanied by composable, agent-friendly APIs that support structured querying and data-driven workflows.

土壤科学多模态数据表征学习数据融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。