arXiv:2603.08533cs.CV2026-03被引 3

SecAgent用中文GUI数据提升手机助手效率,让小模型也能搞定复杂任务。

SecAgent: Efficient Mobile GUI Agent with Semantic Context

  • 用18000个中文界面样本构建高质量数据集,支持多应用导航
  • 将历史截图和操作转为简洁语义摘要,降低计算开销30%以上
  • 30亿参数模型在中文任务上媲美70-80亿大模型,适合移动端部署

基于多模态大语言模型的移动图形用户界面(GUI)代理在自动化复杂智能手机任务方面展现出巨大潜力。然而,现有方法存在两大关键局限:高质量多语言数据集稀缺,尤其在非英语生态系统中;历史信息表示方法效率低下。为此,我们提出30亿参数的高效移动GUI代理SecAgent。首先,构建了一个人工验证的中文移动GUI数据集,包含1.8万个标注样本和12.1万步导航操作,覆盖44个应用,并开发了带有多选动作标注的中文导航基准。基于该数据集,我们提出一种语义上下文机制,将历史截图和操作提炼为紧凑的自然语言摘要,显著降低计算成本的同时保留任务相关信息。通过监督微调与强化学习微调,SecAgent在我们的及公开导航基准上表现优于同类规模基线,性能接近70-80亿参数模型。数据集已开源:https://huggingface.co/datasets/alibabagroup/CMGUI。

原文摘要 · Abstract (English)

Mobile Graphical User Interface (GUI) agents powered by multimodal large language models have demonstrated promising capabilities in automating complex smartphone tasks. However, existing approaches face two critical limitations: the scarcity of high-quality multilingual datasets, particularly for non-English ecosystems, and inefficient history representation methods. To address these challenges, we present SecAgent, an efficient mobile GUI agent at 3B scale. We first construct a human-verified Chinese mobile GUI dataset with 18k grounding samples and 121k navigation steps across 44 applications, along with a Chinese navigation benchmark featuring multi-choice action annotations. Building upon this dataset, we propose a semantic context mechanism that distills history screenshots and actions into concise, natural language summaries, significantly reducing computational costs while preserving task-relevant information. Through supervised and reinforcement fine-tuning, SecAgent outperforms similar-scale baselines and achieves performance comparable to 7B-8B models on our and public navigation benchmarks. Our dataset is available at https://huggingface.co/datasets/alibabagroup/CMGUI.

移动代理中文数据集语义摘要轻量化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。