用视觉+大模型让机器人帮老人操作手机,成功率超81%。
PeriGuru: A Peripheral Robotic Mobile App Operation Assistant based on GUI Image Understanding and Prompting with LLM
- 通过图像识别和大模型理解手机界面,生成操作指令
- 在测试集上成功率达81.94%,是无视觉解析方法的两倍以上
- 适合助老、残障人士辅助使用智能手机,隐私安全
智能手机显著提升了人们的日常学习、沟通与娱乐体验,已成为现代生活不可或缺的部分。然而,老年人及残障人士在使用智能手机时面临诸多困难,亟需移动应用操作助手(即移动应用代理)。为应对隐私、权限及跨平台兼容性问题,本文提出并开发了PeriGuru——一种基于GUI图像理解与大语言模型(LLM)提示的外设式机器人手机操作助手。PeriGuru利用一系列计算机视觉技术分析手机界面截图,并通过大语言模型做出操作决策,最终由机械臂执行。在测试任务集上,PeriGuru的成功率达到81.94%,较无GUI图像解析与提示设计的方法提升逾一倍。代码已开源:https://github.com/Z2sJ4t/PeriGuru。
原文摘要 · Abstract (English)
Smartphones have significantly enhanced our daily learning, communication, and entertainment, becoming an essential component of modern life. However, certain populations, including the elderly and individuals with disabilities, encounter challenges in utilizing smartphones, thus necessitating mobile app operation assistants, a.k.a. mobile app agent. With considerations for privacy, permissions, and cross-platform compatibility issues, we endeavor to devise and develop PeriGuru in this work, a peripheral robotic mobile app operation assistant based on GUI image understanding and prompting with Large Language Model (LLM). PeriGuru leverages a suite of computer vision techniques to analyze GUI screenshot images and employs LLM to inform action decisions, which are then executed by robotic arms. PeriGuru achieves a success rate of 81.94% on the test task set, which surpasses by more than double the method without PeriGuru's GUI image interpreting and prompting design. Our code is available on https://github.com/Z2sJ4t/PeriGuru.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。