通过迭代偏好学习提升移动端智能体的思维能力
MobileIPL: Enhancing Mobile Agents Thinking Process via Iterative Preference Learning
- 构建CoaT树并用规则奖励评分,反向优化思维步骤
- 在三个基准上超越OS-ATLAS和UI-TARS等强基线
- 适合需要高泛化能力的移动端自动化任务研究者
基于视觉语言模型的移动端智能体在GUI任务中表现优异,但其思维轨迹(CoaT)数据稀缺,限制了表达力与泛化能力。现有自训练方法或忽略中间推理正确性,或依赖昂贵的过程级标注构建过程奖励模型(PRM)。为此,我们提出迭代偏好学习(IPL),通过迭代采样构建CoaT树,以规则奖励评分叶节点,并反向传播反馈生成思维级直接偏好优化(T-DPO)对。为防止预热阶段过拟合,引入三阶段指令演化:利用GPT-4o根据真实移动端UI截图生成多样化问答对,提升泛化性与布局理解能力。在三个标准移动端GUI智能体基准上的实验表明,MobileIPL优于强基线,包括持续预训练模型OS-ATLAS和UI-TARS,达到当前最佳性能,并在跨域场景中表现出强泛化能力。
原文摘要 · Abstract (English)
The Chain of Action-Planning Thoughts (CoaT) paradigm has been shown to improve the reasoning performance of VLM-based mobile agents in GUI tasks. However, the scarcity of diverse CoaT trajectories limits the expressiveness and generalization ability of such agents. While self-training is commonly employed to address data scarcity, existing approaches either overlook the correctness of intermediate reasoning steps or depend on expensive process-level annotations to construct process reward models (PRM). To address the above problems, we propose an Iterative Preference Learning (IPL) that constructs a CoaT-tree through interative sampling, scores leaf nodes using rule-based reward, and backpropagates feedback to derive Thinking-level Direct Preference Optimization (T-DPO) pairs. To prevent overfitting during warm-up supervised fine-tuning, we further introduce a three-stage instruction evolution, which leverages GPT-4o to generate diverse Q\&A pairs based on real mobile UI screenshots, enhancing both generality and layout understanding. Experiments on three standard Mobile GUI-agent benchmarks demonstrate that our agent MobileIPL outperforms strong baselines, including continual pretraining models such as OS-ATLAS and UI-TARS. It achieves state-of-the-art performance across three standard Mobile GUI-Agents benchmarks and shows strong generalization to out-of-domain scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。