arXiv:2505.19095cs.AI2025-05被引 2

让AI在开放图形界面中自主探索,不依赖人工数据集。

ScreenExplorer: Training a Vision-Language Model for Diverse Exploration in Open GUI World

  • 用动态环境训练视觉语言模型,提升泛化能力。
  • 引入好奇心奖励机制,解决初期探索难问题。
  • 适合研究通用智能体与交互式系统的人参考。

大型语言模型(LLMs)的快速发展激发了在图形用户界面(GUI)环境中构建通用人工智能(AGI)的兴趣。然而,现有基于LLM或视觉语言模型(VLM)的GUI智能体往往难以泛化到新环境,且严重依赖手工标注的多样化数据集。为克服这些局限,我们提出ScreenExplorer,一种通过群体相对策略优化(GRPO)在真实、动态、开放的GUI环境中训练的VLM。创新性地,我们引入基于世界模型的趣味性奖励函数,帮助智能体克服探索的冷启动阶段。此外,经验流蒸馏进一步增强了模型的探索能力。我们的训练框架显著提升了模型在开放GUI环境中的探索性能,训练后的模型展现出更强的环境适应性和持续探索能力,优于静态部署模型。研究成果为复杂交互场景下具备自我改进能力的AGI系统提供了可扩展路径。

原文摘要 · Abstract (English)

The rapid progress of large language models (LLMs) has sparked growing interest in building Artificial General Intelligence (AGI) within Graphical User Interface (GUI) environments. However, existing GUI agents based on LLMs or vision-language models (VLMs) often fail to generalize to novel environments and rely heavily on manually curated, diverse datasets. To overcome these limitations, we introduce ScreenExplorer, a VLM trained via Group Relative Policy Optimization(GRPO) in real, dynamic, and open-ended GUI environments. Innovatively, we introduced a world-model-based curiosity reward function to help the agent overcome the cold-start phase of exploration. Additionally, distilling experience streams further enhances the model's exploration capabilities. Our training framework enhances model exploration in open GUI environments, with trained models showing better environmental adaptation and sustained exploration compared to static deployment models. Our findings offer a scalable pathway toward AGI systems with self-improving capabilities in complex interactive settings.

视觉语言模型强化学习人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。