用结构化预测提升移动端GUI代理的每步操作准确性
StepReflect: Structured UI Transition Reflection for Mobile GUI Agents

- 将每步GUI操作反思建模为条件结构化预测,依赖显式状态转移规范和视觉证据
- 在AndroidWorld上达82.16%的转换级准确率,比零样本GPT-5.2高11.83个百分点
- 可本地部署,降低API成本,适合需要长期可靠执行的移动端自动化场景
自主移动GUI代理需要精确的动作反思以保障长时程执行的可靠性。现有方法依赖每次动作后的开放式多模态推理,成本高且与GUI状态转移的结构性不匹配。本文提出StepReflect,将每步GUI反思建模为基于显式转移规范和配对视觉证据的监督结构化预测。模型通过分阶段训练流程(监督微调、师生蒸馏、偏好与奖励精调)获得。离线测试中,8B规模模型在AndroidWorld上实现82.16%的转换级准确率,较相同输入条件下零样本GPT-5.2高出11.83个百分点。在线测试中,在M3A、Agent-SAMA、MAI-UI-8B和Seed-2.0-Pro四个代理配置下,三项任务成功率更高,第四项仅差一次成功任务;同时在全部配置中相比GPT基反思显著降低付费API开销。结果表明,StepReflect是长时程移动GUI代理中可本地部署的高效替代方案。
原文摘要 · Abstract (English)
Autonomous mobile GUI agents require accurate action reflection for reliable long-horizon execution. Existing approaches rely on open-ended multimodal reasoning after each action, which is costly and poorly matched to the structured nature of GUI state transitions. We propose StepReflect, which formulates per-step GUI reflection as supervised structured prediction conditioned on explicit transition specifications and paired visual evidence. StepReflect is trained through a staged pipeline combining supervised fine-tuning, teacher-student distillation, and preference- and reward-based refinement. Offline, the resulting 8B model achieves 82.16% transition-level accuracy on AndroidWorld, exceeding zero-shot GPT-5.2 by 11.83 percentage points under the same structured input. Online, across M3A, Agent-SAMA, MAI-UI-8B, and Seed-2.0-Pro, StepReflect achieves higher task success in three of four agent configurations and remains within one successful task of the GPT-5.2 Reflection Agent in the fourth. It also reduces paid API charges relative to GPT-based reflection in all four configurations. These results establish StepReflect as a practical, locally deployable alternative to repeated frontier-model reflection for long-horizon mobile GUI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。