分层视觉语言代理让手机操控更智能,成功率超87%。
Hi-Agent: Hierarchical Vision-Language Agents for Mobile Device Control
- 分高阶推理与低阶动作两层,联合优化提升决策能力。
- 在AitW基准上达87.9%成功率,远超现有方法。
- 无需额外标注,零样本泛化能力强,适合复杂场景。
构建能自主操作移动设备的智能体受到广泛关注。尽管视觉语言模型(VLMs)展现出潜力,但多数方法依赖直接的状态到动作映射,缺乏结构化推理与规划,因此在新任务或未见过的界面布局下泛化能力差。我们提出Hi-Agent,一种可训练的分层视觉语言代理,包含高阶推理模型与低阶动作模型,并实现联合优化。为提升训练效率,我们将多步决策重构为一系列单步子目标,并提出前瞻优势函数,利用低阶模型的执行反馈指导高阶优化。该设计缓解了长程任务中组相对策略优化(GRPO)的路径爆炸问题,实现稳定、无批评器的联合训练。Hi-Agent在Android-in-the-Wild(AitW)基准上取得87.9%的任务成功率,显著优于三种范式:提示驱动(AppAgent: 17.7%)、监督学习(Filtered BC: 54.5%)和强化学习(DigiRL: 71.9%)。其在ScreenSpot-v2基准上也展现良好零样本泛化能力。在更具挑战性的AndroidWorld基准上,随着骨干网络增大,性能持续提升,表现出强适应性。
原文摘要 · Abstract (English)
Building agents that autonomously operate mobile devices has attracted increasing attention. While Vision-Language Models (VLMs) show promise, most existing approaches rely on direct state-to-action mappings, which lack structured reasoning and planning, and thus generalize poorly to novel tasks or unseen UI layouts. We introduce Hi-Agent, a trainable hierarchical vision-language agent for mobile control, featuring a high-level reasoning model and a low-level action model that are jointly optimized. For efficient training, we reformulate multi-step decision-making as a sequence of single-step subgoals and propose a foresight advantage function, which leverages execution feedback from the low-level model to guide high-level optimization. This design alleviates the path explosion issue encountered by Group Relative Policy Optimization (GRPO) in long-horizon tasks and enables stable, critic-free joint training. Hi-Agent achieves a new State-Of-The-Art (SOTA) 87.9% task success rate on the Android-in-the-Wild (AitW) benchmark, significantly outperforming prior methods across three paradigms: prompt-based (AppAgent: 17.7%), supervised (Filtered BC: 54.5%), and reinforcement learning-based (DigiRL: 71.9%). It also demonstrates competitive zero-shot generalization on the ScreenSpot-v2 benchmark. On the more challenging AndroidWorld benchmark, Hi-Agent also scales effectively with larger backbones, showing strong adaptability in high-complexity mobile control scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。