将多智能体系统升级为可自组织、自改进的AI公司,实现动态协作与持续优化。
From Skills to Talent: Organising Heterogeneous Agents as a Real-World Company

- 用‘人才’封装智能体能力,通过市场机制按需招募和重组团队。
- 在PRDBench上达成84.67%成功率,比现有方法高15.48个百分点。
- 适合需要跨领域灵活应对复杂任务的AI系统研发者使用。
个体智能体的能力因模块化技能与工具集成迅速提升,但多智能体系统仍受限于固定团队结构、强耦合协调逻辑和会话绑定学习。本文指出其深层缺失:缺乏一个独立于个体知识的组织层来管理智能体的组建、治理与进化。为此,提出OneManCompany(OMC)框架,将技能、工具和运行配置封装为可迁移的“人才”身份,通过类型化组织接口抽象异构后端。社区驱动的“人才市场”支持动态招聘,实时填补能力缺口并重构组织。组织决策采用探索-执行-评审(E²R)树搜索机制,统一规划、执行与评估,在层级闭环中分解任务并聚合结果,实现系统性反思与优化。该机制保证终止性和无死锁,类比人类企业反馈流程。实验表明,OMC在PRDBench上成功率达84.67%,超越当前最优水平15.48个百分点;跨领域案例进一步验证其通用性。
原文摘要 · Abstract (English)
Individual agent capabilities have advanced rapidly through modular skills and tool integrations, yet multi-agent systems remain constrained by fixed team structures, tightly coupled coordination logic, and session-bound learning. We argue that this reflects a deeper absence: a principled organisational layer that governs how a workforce of agents is assembled, governed, and improved over time, decoupled from what individual agents know. To fill this gap, we introduce \emph{OneManCompany (OMC)}, a framework that elevates multi-agent systems to the organisational level. OMC encapsulates skills, tools, and runtime configurations into portable agent identities called \emph{Talents}, orchestrated through typed organisational interfaces that abstract over heterogeneous backends. A community-driven \emph{Talent Market} enables on-demand recruitment, allowing the organisation to close capability gaps and reconfigure itself dynamically during execution. Organisational decision-making is operationalised through an \emph{Explore-Execute-Review} ($\text{E}^2$R) tree search, which unifies planning, execution, and evaluation in a single hierarchical loop: tasks are decomposed top-down into accountable units and execution outcomes are aggregated bottom-up to drive systematic review and refinement. This loop provides formal guarantees on termination and deadlock freedom while mirroring the feedback mechanisms of human enterprises. Together, these contributions transform multi-agent systems from static, pre-configured pipelines into self-organising and self-improving AI organisations capable of adapting to open-ended tasks across diverse domains. Empirical evaluation on PRDBench shows that OMC achieves an $84.67\%$ success rate, surpassing the state of the art by $15.48$ percentage points, with cross-domain case studies further demonstrating its generality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。