arXiv:2505.10887cs.AI2025-05NeurIPS被引 7

多模态通用智能体实现自动化电脑操作,支持文本图像音频视频协同交互。

InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction

  • 模块化架构融合工具型与纯视觉代理,分步协作解决独立任务。
  • 在OSWorld上达7.27%准确率,优于Claude-Computer-Use。
  • 适用于复杂工具调用场景,适合研究自动化交互的开发者。

本文提出 extsc{InfantAgent-Next},一种能够以多模态方式(包括文本、图像、音频和视频)与计算机交互的通用智能体。与现有方法仅依赖单一大模型或仅提供工作流模块化不同,该智能体在高度模块化的架构中集成工具型代理与纯视觉代理,使不同模型可分步协作处理解耦任务。其通用性通过在纯视觉现实基准(如OSWorld)以及更通用或工具密集型基准(如GAIA和SWE-Bench)上的评估得到验证。具体而言,在OSWorld上取得7.27%的准确率,高于Claude-Computer-Use。代码与评估脚本已开源至https://github.com/bin123apple/InfantAgent。

原文摘要 · Abstract (English)

This paper introduces \textsc{InfantAgent-Next}, a generalist agent capable of interacting with computers in a multimodal manner, encompassing text, images, audio, and video. Unlike existing approaches that either build intricate workflows around a single large model or only provide workflow modularity, our agent integrates tool-based and pure vision agents within a highly modular architecture, enabling different models to collaboratively solve decoupled tasks in a step-by-step manner. Our generality is demonstrated by our ability to evaluate not only pure vision-based real-world benchmarks (i.e., OSWorld), but also more general or tool-intensive benchmarks (e.g., GAIA and SWE-Bench). Specifically, we achieve $\mathbf{7.27\%}$ accuracy on OSWorld, higher than Claude-Computer-Use. Codes and evaluation scripts are open-sourced at https://github.com/bin123apple/InfantAgent.

多模态智能体自动化通用代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。