arXiv:2602.02559cs.AIcs.CV2026-02被引 2

无需训练,通过交互让智能体自动掌握地球观测技能

Experience-Driven Multi-Agent Systems Are Training-free Context-aware Earth Observers

  • 用多智能体架构拆解任务,逐层试错工具参数配置
  • 在三个地球观测基准上平均提升12%成功率,跨大模型通用
  • 适合需要精细操作和容错的复杂遥感分析场景

近期进展使大语言模型(LLM)代理能够通过协调外部工具解决复杂任务。然而,在需长时程执行、多模态紧密协作且严格遵守隐式工具约束的专用领域中,这些代理仍表现不佳。地球观测(EO)任务正是此类挑战的典型:其涉及多模态、多时相数据输入,以及地理知识约束(如光谱库、空间推理等),许多高层计划会因细微执行错误在流程中累积并导致最终结果失效。核心难点在于现有代理缺乏从交互中学习细粒度工具级专业知识的机制。若无此能力,它们无法可靠配置工具参数或在执行中恢复,限制了在复杂EO工作流中的有效性。为此,我们提出 extbf{GeoEvolver},一种自演化的多智能体系统(MAS),使LLM代理能在不更新参数的前提下,通过结构化交互获取地球观测专业知识。GeoEvolver利用检索增强的多智能体编排器将每个查询分解为独立子目标,然后在子目标层面探索多样化的工具参数配置。成功模式与失败根因分析被提炼至动态记忆库,为后续查询提供上下文示范。在三个集成工具的地球观测基准上的实验表明,GeoEvolver持续提升了端到端任务成功率,跨多个LLM骨干网络平均提升12%,证明了地球观测专长可从高效、细粒度的环境交互中逐步涌现。

原文摘要 · Abstract (English)

Recent advances have enabled large language model (LLM) agents to solve complex tasks by orchestrating external tools. However, these agents often struggle in specialized, tool-intensive domains that demand long-horizon execution, tight coordination across modalities, and strict adherence to implicit tool constraints. Earth Observation (EO) tasks exemplify this challenge due to the multi-modal and multi-temporal data inputs, as well as the requirements of geo-knowledge constraints (spectrum library, spatial reasoning, etc): many high-level plans can be derailed by subtle execution errors that propagate through a pipeline and invalidate final results. A core difficulty is that existing agents lack a mechanism to learn fine-grained, tool-level expertise from interaction. Without such expertise, they cannot reliably configure tool parameters or recover from mid-execution failures, limiting their effectiveness in complex EO workflows. To address this, we introduce \textbf{GeoEvolver}, a self-evolving multi-agent system~(MAS) that enables LLM agents to acquire EO expertise through structured interaction without any parameter updates. GeoEvolver decomposes each query into independent sub-goals via a retrieval-augmented multi-agent orchestrator, then explores diverse tool-parameter configurations at the sub-goal level. Successful patterns and root-cause attribution from failures are then distilled in an evolving memory bank that provides in-context demonstrations for future queries. Experiments on three tool-integrated EO benchmarks show that GeoEvolver consistently improves end-to-end task success, with an average gain of 12\% across multiple LLM backbones, demonstrating that EO expertise can emerge progressively from efficient, fine-grained interactions with the environment.

多智能体地球观测自演化零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。