通过视觉状态转换预训练,提升GUI智能体的泛化能力
Scaling GUI Agents with Visual State Transitions

- 用视觉状态变化联合优化前后向动态模型
- 在多个基准上超越纯轨迹微调基线模型
- 数据越多效果越好,适合跨平台GUI任务
我们提出状态转换预训练(STP),作为GUI智能体的新规模扩展路径。在STP阶段,通过联合优化逆动力学(从状态变化预测动作)和前向动力学(从当前状态与动作预测下一状态),持续对统一多模态模型进行视觉状态转换的预训练。该过程使模型获得更优的动作感知视觉表征和内部的GUI动态世界模型。在后续针对带任务指令轨迹的微调中,我们的STP训练模型在桌面与移动端的多个基准(AgentNetBench、AndroidControl、GUIOdyssey)上均持续优于仅通过直接轨迹微调训练的基线模型。实证研究进一步表明,联合动力学优化相比单目标训练带来稳定提升,且下游性能随转换数据量增加而稳步上升。
原文摘要 · Abstract (English)
We introduce State Transition Pretraining (STP) as a new scaling axis for GUI agents. During the STP stage, we continually pretrain a unified multimodal model on visual state transitions by jointly optimizing inverse dynamics (predicting actions from state changes) and forward dynamics (predicting next states from current states and actions). This optimization equips the model with better action-grounded visual representations and an internal world model of GUI dynamics. When subsequently fine-tuned on trajectories with task instructions, our STP-trained models consistently outperform baselines trained solely via direct trajectory fine-tuning across agent benchmarks in both desktop and mobile GUI scenarios (AgentNetBench, AndroidControl, and GUIOdyssey). Further empirical studies show that joint dynamics optimization yields stable improvements over single-objective training, and downstream performance scales steadily with the volume of transition data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。