arXiv:2512.00846cs.CV2025-12中稿 · WACV 2026 Conferen…被引 2

轻量级GUI自动化模型,通过自适应特征重归一化提升识别精度

AFRAgent : An Adaptive Feature Renormalization Based High Resolution Aware GUI agent

  • 基于BLIP的多模态架构,用自适应归一化增强低分辨率图像特征
  • 在Meta-GUI和AITW上达新SOTA,模型大小不足同类四分之一
  • 适合移动端自动化任务,尤其对算力有限场景友好

移动用户界面自动化需求日益增长,视觉语言模型(VLMs)推动其从生成人类可读指令转向自主执行任务。现有方法虽能直接处理屏幕内容、不依赖设备API,并利用真实上下文理解任务,但受限于视觉编码器特征的空间信息不足,常出现控件识别不准、动作判断错误的问题。同时,高性能模型普遍体积庞大,训练成本高且推理延迟显著。本文提出AFRAgent,一种基于instruct-BLIP的轻量级多模态架构,在保持高性能的同时,模型规模不足最接近的竞争对手的1/4。为增强大语言模型流水线中的图像嵌入,提出自适应特征重归一化(a token-level affine transformation)技术,有效丰富低分辨率图像特征并融合高分辨率细节。在Meta-GUI与AITW基准测试中,AFRAgent建立新的性能标杆,适用于智能手机自动化任务。

原文摘要 · Abstract (English)

There is a growing demand for mobile user interface (UI) automation, driven by its broad applications across industries. With the advent of visual language models (VLMs), GUI automation has progressed from generating text-based instructions for humans to autonomously executing tasks, thus optimizing automation workflows. Recent approaches leverage VLMs for this problem due to their ability to 1) process on-screen content directly, 2) remain independent of device-specific APIs by utilizing human actions (e.g., clicks, typing), and 3) apply real-world contextual knowledge for task understanding. However, these models often have trouble accurately identifying widgets and determining actions due to limited spatial information in vision encoder features. Additionally, top-performing models are often large, requiring extensive training and resulting in inference delays. In this work, we introduce AFRAgent, an instruct-BLIP-based multimodal architecture that achieves superior performance in GUI automation while being less than one-fourth the size of its nearest competitor. To enhance image embeddings in the large language model (LLM) pipeline, we propose an adaptive feature renormalization-based (a token-level affine transformation) technique that effectively enriches low-resolution image embeddings and fuses high-resolution details. We evaluate AFRAgent on Meta-GUI and AITW benchmarks, establishing a new state-of-the-art baseline for smartphone automation.

GUI自动化多模态轻量模型视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。