用多模态大模型自动生成多智能体协作约束,提升自动化水平。
Distributed Multi-Agent Coordination Using Multi-Modal Foundation Models
- 利用视觉与语言指令自动构建多智能体约束
- 提出从神经符号到全神经的智能体范式谱系
- 在三个新任务上验证了不同范式的优劣
分布式约束优化问题(DCOPs)为多智能体协作提供了强大框架,但通常依赖繁琐的手动建模。为此,我们提出VL-DCOPs框架,利用大规模多模态基础模型(LFMs)从视觉和语言指令中自动生成约束。我们设计了一组智能体原型:从将部分算法决策交给LFM的神经符号型,到完全依赖LFM进行协调的全神经型。我们在三个新型VL-DCOP任务上,使用最先进的大语言模型(LLMs)和视觉语言模型(VLMs)评估这些智能体原型,并比较其各自优势与局限。最后,探讨该工作对DCOP文献中更广泛前沿挑战的延伸意义。
原文摘要 · Abstract (English)
Distributed Constraint Optimization Problems (DCOPs) offer a powerful framework for multi-agent coordination but often rely on labor-intensive, manual problem construction. To address this, we introduce VL-DCOPs, a framework that takes advantage of large multimodal foundation models (LFMs) to automatically generate constraints from both visual and linguistic instructions. We then introduce a spectrum of agent archetypes for solving VL-DCOPs: from a neuro-symbolic agent that delegates some of the algorithmic decisions to an LFM, to a fully neural agent that depends entirely on an LFM for coordination. We evaluate these agent archetypes using state-of-the-art LLMs (large language models) and VLMs (vision language models) on three novel VL-DCOP tasks and compare their respective advantages and drawbacks. Lastly, we discuss how this work extends to broader frontier challenges in the DCOP literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。