arXiv:2508.04037cs.AI2025-08被引 7

7B参数模型实现高效计算机操作,性能媲美更大模型。

Evolving in Tasks: Empowering the Multi-modality Large Language Model as the Computer Use Agent

  • 自动生成可验证任务轨迹,提升训练数据质量。
  • 采用分步强化学习,降低长时序任务训练开销。
  • 融合定位与规划能力,无需额外训练即提升表现。

计算机操作代理是人工智能领域的新兴方向,旨在自主执行用户任务,受到产业界和学术界的广泛关注。然而,现有代理的性能仍不足以满足实际部署需求。本文提出自演化代理(SEA)用于计算机操作,并在数据生成、强化学习和模型增强三方面提出核心创新。首先,设计自动化流水线生成可验证的任务轨迹以用于训练;其次,提出高效的分步强化学习方法,显著降低长时序训练的计算开销;最后,引入一种模型增强方法,将定位与规划能力整合至单一模型中,无需额外训练即可提升性能。基于这些创新,我们的SEA(仅70亿参数)在计算机操作任务上的表现超越同规模模型,并达到320亿/720亿参数模型的水平。未来计划开源模型权重及配套代码。

原文摘要 · Abstract (English)

Computer use agents represent an emerging area in artificial intelligence, aiming to operate computers autonomously to fulfill user tasks, attracting significant attention from both industry and academia. However, the performance of existing agents remains insufficient for practical deployment. In this paper, we propose the Self-Evolution Agent (SEA) for computer operation, alongside three core innovations in data generation, reinforcement learning, and model enhancement to develop this agent. Specifically, we first design an automatic pipeline to generate verifiable task trajectories for training. Second, we propose Efficient Step-wise Reinforcement Learning to reduce the substantial computational overhead of long-horizon training. Finally, we introduce a model enhancement method that integrates grounding and planning capabilities into a single model without additional training. Leveraging these innovations, our SEA (with only 7B parameters) outperforms existing models of the same parameter scale and achieves performance comparable to larger models (e.g., 32B/72B parameters) on computer use tasks. We plan to release the model weights and related code as open-source resources in the future.

多模态大模型智能代理强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。