arXiv:2602.21351cs.AIcs.IR2026-02被引 5

用多智能体系统自动挖掘地球科学数据,提升低利用率数据的重用率。

A Hierarchical Multi-Agent System for Autonomous Discovery in Geoscientific Data Archives

  • 分层架构下,智能体按数据类型精准路由并自主纠错。
  • 在海洋学与生态学案例中实现多步流程自动化,错误自修复成功率高。
  • 适合数据科学家、科研人员快速探索异构地质数据集。

地球科学数据的快速积累带来了显著的可扩展性挑战;尽管像PANGAEA这样的数据仓库拥有海量数据集,但引用指标显示大量数据仍被闲置,限制了数据的再利用。本文提出PANGAEA-GPT,一种用于自主数据发现与分析的分层多智能体框架。不同于常规的大语言模型封装,该架构采用集中式监督-工作模式,具备数据类型感知的路由机制、沙箱内确定性代码执行以及基于执行反馈的自我修正能力,使智能体能够诊断并解决运行时错误。通过物理海洋学和生态学的应用场景,验证了系统在极少人工干预下执行复杂多步工作流的能力。该框架为查询和分析异构数据仓库中的数据提供了一种协同智能体工作流的方法。

原文摘要 · Abstract (English)

The rapid accumulation of Earth science data has created a significant scalability challenge; while repositories like PANGAEA host vast collections of datasets, citation metrics indicate that a substantial portion remains underutilized, limiting data reusability. Here we present PANGAEA-GPT, a hierarchical multi-agent framework designed for autonomous data discovery and analysis. Unlike standard Large Language Model (LLM) wrappers, our architecture implements a centralized Supervisor-Worker topology with strict data-type-aware routing, sandboxed deterministic code execution, and self-correction via execution feedback, enabling agents to diagnose and resolve runtime errors. Through use-case scenarios spanning physical oceanography and ecology, we demonstrate the system's capacity to execute complex, multi-step workflows with minimal human intervention. This framework provides a methodology for querying and analyzing heterogeneous repository data through coordinated agent workflows.

多智能体数据挖掘地球科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。