用精选数据和轻量训练,让小模型也能高效做界面推理
An Efficient Training Pipeline for Reasoning Graphical User Interface Agents
- 从480万合成数据中筛选出1.2万条高质量、多样化的样本
- 30亿参数模型在筛选数据上训练后性能超过更大模型
- 适合想打造轻量级智能界面助手的研究者和开发者
视觉定位是将自然语言查询与图像区域对应的关键任务,对具备推理能力的图形用户界面代理至关重要。现有方法多依赖大规模且嘈杂的合成数据集。本文提出一种高效的训练流程,结合基于模型的数据过滤与参数高效微调。从480万条合成样本中,先识别困难样本,剔除错位样本,再选取1.2万条具有多样性的跨模态实例。在此数据集上,使用30亿参数的视觉-语言模型在三种训练模式下进行训练:监督微调、思维链增强微调以及基于组相对策略优化的强化学习。使用过滤后数据和轻量化策略训练的模型,在ScreenSpot、Multimodal-Mind2Web和AndroidControl等基准测试中达到或超越更大规模基线模型的表现。结果表明,系统性数据筛选与稳健适配可媲美大规模训练,使小型但强大的多模态推理代理成为可能。
原文摘要 · Abstract (English)
Visual grounding is the task of localising image regions from natural language queries and is critical for reasoning capable Graphical User Interface agents. Many existing methods rely on massive, noisy synthetic datasets. This work introduces an efficient training pipeline that combines model-based data filtering with parameter-efficient fine-tuning. From 4.8M synthetic examples, 12K clean and diverse instances are curated by first identifying challenging cases, removing misaligned and then selecting a diverse set of multimodal instances. On this data, a 3B-parameter Vision-Language Model is trained under three regimes: supervised fine-tuning, chain-of-thought-augmented fine-tuning, and reinforcement learning via Group Relative Policy Optimization. Models trained with the filtered data and lightweight training strategies match or surpass larger baselines on benchmarks such as ScreenSpot, Multimodal-Mind2Web, and AndroidControl. These results demonstrate that principled data curation and robust adaptation can rival large-scale training, enabling compact yet capable multimodal reasoning agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。