arXiv:2604.01647cs.AI2026-04中稿 · PEARC 2026被引 1

用多智能体系统提升环境数据管理的可靠性,防止错误数据发布。

Exploring Robust Multi-Agent Workflows for Environmental Data Management

论文配图:Exploring Robust Multi-Agent Workflows for Environmental Data Management
图 1 · 摘自论文原文
  • 分三轨知识架构+角色分离设计,让智能体行为可追溯、可验证。
  • 在真实数据集上实现两天完成任务,错误在发布前被自动拦截。
  • 适合需要高可靠性的科研数据管理团队,尤其是长期监测项目。

将大模型驱动的智能体融入环境数据的FAIR管理极具吸引力——它们能外化操作知识,跨异构数据和动态规范扩展数据整理能力。但用概率性流程替代确定性组件会改变故障模式:大模型生成看似合理却错误的输出,可能通过表面检查并导致不可逆操作,如DOI注册与公开发布。本文提出EnviSmart,一个部署于校园级存储基础设施上的生产级数据管理系统。该系统通过两项机制将可靠性作为架构属性:三轨知识架构,将治理约束、领域知识(可检索上下文)和技能(工具使用流程)作为持久、相互关联的实体外化;角色分离的多智能体设计,通过确定性验证器与审计交接,在信任边界恢复“失败即停止”语义,防止不可逆步骤。在两个生产部署中对比验证:大学地理信息中心生态档案库(849个已整理数据集)作为单智能体基线;SF2Bench复合洪水基准测试包含2,452个监测站、8,557份已发布文件,跨度39年,用于验证多智能体流程。多智能体方案提升了效率——单操作员两天内完成,且跨部署重复利用成果;同时提升可靠性:审计交接提前发现影响全部2,452个站点的坐标转换错误,阻止其发布。代表性事件(ISS-004)显示边界式隔离,检测延迟仅10分钟,零用户暴露,80分钟内解决。

原文摘要 · Abstract (English)

Embedding LLM-driven agents into environmental FAIR data management is compelling - they can externalize operational knowledge and scale curation across heterogeneous data and evolving conventions. However, replacing deterministic components with probabilistic workflows changes the failure mode: LLM pipelines may generate plausible but incorrect outputs that pass superficial checks and propagate into irreversible actions such as DOI minting and public release. We introduce EnviSmart, a production data management system deployed on campus-wide storage infrastructure for environmental research. EnviSmart treats reliability as an architectural property through two mechanisms: a three-track knowledge architecture that externalizes behaviors (governance constraints), domain knowledge (retrievable context), and skills (tool-using procedures) as persistent, interlocking artifacts; and a role-separated multi-agent design where deterministic validators and audited handoffs restore fail-stop semantics at trust boundaries before irreversible steps. We compare two production deployments. The University's GIS Center Ecological Archive (849 curated datasets) serves as a single-agent baseline. SF2Bench, a compound flooding benchmark comprising 2,452 monitoring stations and 8,557 published files spanning 39 years, validates the multi-agent workflow. The multi-agent approach improved both efficiency - completed by a single operator in two days with repeated artifact reuse across deployments - and reliability: audited handoffs detected and blocked a coordinate transformation error affecting all 2,452 stations before publication. A representative incident (ISS-004) demonstrated boundary-based containment with 10-minute detection latency, zero user exposure, and 80-minute resolution. This paper has been accepted at PEARC 2026.

多智能体数据管理环境科学可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。