arXiv:2509.17336cs.MMcs.CL2025-09被引 13

Mano通过多阶段训练提升GUI操作成功率,解决视觉语言模型在界面交互中的局限性。

Mano Technical Report

  • 用模拟环境生成高保真数据,结合三阶段训练提升决策能力
  • 在Mind2Web和OSWorld上达成最优表现,成功率显著提升
  • 适合需要高可靠性的自动化测试与智能助手场景

图形用户界面(GUI)是人机交互的主要方式,但自动化界面操作仍面临视觉元素复杂、环境动态变化及多步推理需求等挑战。现有基于视觉语言模型(VLM)的方法常受限于分辨率不足、领域偏差以及序列决策能力弱。为此,我们提出Mano——一个基于大规模网页与系统数据预训练的多模态基础模型构建的鲁棒GUI代理。该方法融合新型模拟环境实现高保真数据生成,采用三阶段训练流程(监督微调、离线强化学习、在线强化学习),并引入验证模块实现错误恢复。Mano在多个GUI基准测试中表现卓越,包括Mind2Web和OSWorld,成功率达显著提升,操作准确率亦大幅优化。本工作揭示了强化学习与VLM有效结合的关键路径,强调领域特定数据、迭代训练与整体奖励设计的重要性。

原文摘要 · Abstract (English)

Graphical user interfaces (GUIs) are the primary medium for human-computer interaction, yet automating GUI interactions remains challenging due to the complexity of visual elements, dynamic environments, and the need for multi-step reasoning. Existing methods based on vision-language models (VLMs) often suffer from limited resolution, domain mismatch, and insufficient sequential decisionmaking capability. To address these issues, we propose Mano, a robust GUI agent built upon a multi-modal foundation model pre-trained on extensive web and computer system data. Our approach integrates a novel simulated environment for high-fidelity data generation, a three-stage training pipeline (supervised fine-tuning, offline reinforcement learning, and online reinforcement learning), and a verification module for error recovery. Mano demonstrates state-of-the-art performance on multiple GUI benchmarks, including Mind2Web and OSWorld, achieving significant improvements in success rate and operational accuracy. Our work provides new insights into the effective integration of reinforcement learning with VLMs for practical GUI agent deployment, highlighting the importance of domain-specific data, iterative training, and holistic reward design.

GUI自动化视觉语言模型强化学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。