让手机智能助手读懂用户个性化指令并自动执行
PerPilot: Personalizing VLM-based Mobile Agents via Memory and Exploration
- 用记忆检索和推理探索双重机制识别个性化指令
- 在多个场景中实现低干预执行且越用越准
- 适合开发能理解用户习惯的下一代智能助手
基于视觉语言模型(VLM)的移动智能体在执行指令驱动任务方面展现出巨大潜力,但普遍难以处理包含模糊、用户特有上下文的个性化指令,这一问题此前未受重视。本文定义了个性化指令,并提出PerInstruct数据集,涵盖多种移动场景下的多样化个性化指令。针对现有移动智能体个性化能力不足的问题,提出PerPilot框架,该框架基于大语言模型(LLM),使智能体能够自主感知、理解并执行个性化指令。PerPilot通过记忆检索与推理探索两种互补方式识别个性化元素并完成任务。实验表明,PerPilot能以最小用户干预有效处理个性化任务,并随着使用持续提升性能,凸显了个性化感知推理对下一代移动智能体的重要性。数据集与代码已公开:https://github.com/xinwang-nwpu/PerPilot。
原文摘要 · Abstract (English)
Vision language model (VLM)-based mobile agents show great potential for assisting users in performing instruction-driven tasks. However, these agents typically struggle with personalized instructions -- those containing ambiguous, user-specific context -- a challenge that has been largely overlooked in previous research. In this paper, we define personalized instructions and introduce PerInstruct, a novel human-annotated dataset covering diverse personalized instructions across various mobile scenarios. Furthermore, given the limited personalization capabilities of existing mobile agents, we propose PerPilot, a plug-and-play framework powered by large language models (LLMs) that enables mobile agents to autonomously perceive, understand, and execute personalized user instructions. PerPilot identifies personalized elements and autonomously completes instructions via two complementary approaches: memory-based retrieval and reasoning-based exploration. Experimental results demonstrate that PerPilot effectively handles personalized tasks with minimal user intervention and progressively improves its performance with continued use, underscoring the importance of personalization-aware reasoning for next-generation mobile agents. The dataset and code are available at: https://github.com/xinwang-nwpu/PerPilot
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。