arXiv:2604.17817cs.HCcs.AI2026-04

测试大模型手机自动化能力,发现看图比看字效果略好

Do LLMs Need to See Everything? A Benchmark and Study of Failures in LLM-driven Smartphone Automation using Screentext vs. Screenshots

论文配图:Do LLMs Need to See Everything? A Benchmark and Study of Failures in LLM-driven Smartphone Automation using Screentext vs. Screenshots
图 1 · 摘自论文原文
  • 用文本和截图两种方式测试大模型完成75个手机任务
  • 多模态输入成功率仅比纯文本高1.8个百分点
  • 总结出界面设计、模型理解等常见失败原因

随着大语言模型(LLMs)的快速发展,移动代理已成为实现手机自动化的有前景工具,能够模拟人类在屏幕上交互以完成复杂任务。然而,这些代理常因准确率低、误解用户指令或在挑战性任务中失败而表现不佳,且此前研究极少深入分析其失败原因。为此,我们提出DailyDroid基准,包含25个Android应用中的75项任务,覆盖五个场景与三个难度等级,贴近日常手机使用。我们在GPT-4o和o4-mini上对文本仅输入与多模态(文本+截图)输入进行了300次试验评估,结果显示两者性能相近,多模态输入仅略微提升成功率(+1.8%)。通过深入失败分析,我们整理出一份常见失败手册。研究揭示了用户界面可访问性、输入模态及大模型/应用设计中的关键问题,为未来移动代理、应用开发和界面设计提供了重要启示。

原文摘要 · Abstract (English)

With the rapid advancement of large language models (LLMs), mobile agents have emerged as promising tools for phone automation, simulating human interactions on screens to accomplish complex tasks. However, these agents often suffer from low accuracy, misinterpretation of user instructions, and failure on challenging tasks, with limited prior work examining why and where they fail. To address this, we introduce DailyDroid, a benchmark of 75 tasks in five scenarios across 25 Android apps, spanning three difficulty levels to mimic everyday smartphone use. We evaluate it using text-only and multimodal (text + screenshot) inputs on GPT-4o and o4-mini across 300 trials, revealing comparable performance with multimodal inputs yielding marginally higher success rates. Through in-depth failure analysis, we compile a handbook of common failures. Our findings reveal critical issues in UI accessibility, input modalities, and LLM/app design, offering implications for future mobile agents, applications, and UI development.

手机自动化多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。