arXiv:2510.07423cs.AI2025-10

让AI通过探索与动态调整计划来解决复杂问题。

ProSEA: Problem Solving via Exploration Agents

  • 分层架构下,管理代理协调专家代理迭代优化解题策略。
  • 在FinanceBench上无需人类干预即超越现有最优基线。
  • 失败时报告原因与新约束,支持自动计划重构。

大型语言模型(LLMs)已使AI代理能够处理日益复杂的任务。然而,大多数现有代理仍受限于静态规划和脆弱的交互,无法实现真正的协作或自适应推理。我们提出ProSEA,一种模块化、通用的多代理框架,通过探索与计划演进实现迭代式问题求解。ProSEA采用分层架构,由管理代理协调领域专用的专家代理,分解任务,并根据失败尝试的结构化反馈自适应重规划。不同于以往系统,ProSEA代理不仅报告成功或失败,还提供失败原因和新发现的约束,从而基于探索轨迹动态优化计划。该框架可自主运行,必要时也可无缝融入人类协作。在具有挑战性的FinanceBench基准上的实验表明,ProSEA即使无须人类反馈,也优于现有最先进基线,在高推理需求任务中表现稳健。这些结果凸显了ProSEA作为更透明、自适应且与人类对齐的AI代理基础的潜力。

原文摘要 · Abstract (English)

Large language models (LLMs) have empowered AI agents to tackle increasingly complex tasks. However, most existing agents remain limited to static planning and brittle interactions, falling short of true collaboration or adaptive reasoning. We introduce ProSEA, a modular, general-purpose multi-agent framework designed for iterative problem solving through exploration and plan evolution. ProSEA features a hierarchical architecture in which a Manager Agent orchestrates domain-specialized Expert Agents, decomposes tasks, and adaptively replans based on structured feedback from failed attempts. Unlike prior systems, ProSEA agents report not only success or failure but also detailed reasons for failure and newly discovered constraints, enabling dynamic plan refinement informed by exploratory traces. The framework operates autonomously but supports seamless integration with human collaborators when needed. Experiments on the challenging FinanceBench benchmark demonstrate that ProSEA, even without human feedback, outperforms state-of-the-art baselines and achieves robust performance across reasoning-heavy tasks. These results underscore ProSEA's potential as a foundation for more transparent, adaptive, and human-aligned AI agents.

多智能体推理增强自主规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。