arXiv:2508.09057cs.CL2025-08被引 7

构建真实场景移动代理评估基准,提升复杂指令理解能力

MVISU-Bench: Benchmarking Mobile Agents for Real-World Tasks by Multi-App, Vague, Interactive, Single-App and Unethical Instructions

  • 设计多任务、模糊、交互式等五类真实指令的评估框架
  • 在137个应用中测试404项任务,性能提升19.55%(含恶意指令+53.52%)
  • 轻量模块Aider可动态澄清意图,适合集成到各类移动代理系统

随着大视觉语言模型在推理与视觉理解上的进展,移动代理正快速兴起以满足用户自动化需求。然而现有评估基准与真实场景脱节,难以覆盖用户多样化复杂需求。基于大量用户问卷调研,我们识别出五类任务:多应用、模糊、交互式、单应用及非道德指令。围绕这些任务,提出双语基准MVISU-Bench,涵盖137个移动应用中的404项任务。同时设计Aider模块,作为即插即用的动态提示器,用于降低风险并澄清用户意图。该模块可轻松集成至多个框架,在MVISU-Bench上相较当前最优表现整体成功率提升19.55%,尤其在非道德指令(+53.52%)和交互式指令(+29.41%)上表现突出。实验揭示了现有移动代理与真实用户期望之间的显著差距。

原文摘要 · Abstract (English)

Given the significant advances in Large Vision Language Models (LVLMs) in reasoning and visual understanding, mobile agents are rapidly emerging to meet users' automation needs. However, existing evaluation benchmarks are disconnected from the real world and fail to adequately address the diverse and complex requirements of users. From our extensive collection of user questionnaire, we identified five tasks: Multi-App, Vague, Interactive, Single-App, and Unethical Instructions. Around these tasks, we present \textbf{MVISU-Bench}, a bilingual benchmark that includes 404 tasks across 137 mobile applications. Furthermore, we propose Aider, a plug-and-play module that acts as a dynamic prompt prompter to mitigate risks and clarify user intent for mobile agents. Our Aider is easy to integrate into several frameworks and has successfully improved overall success rates by 19.55\% compared to the current state-of-the-art (SOTA) on MVISU-Bench. Specifically, it achieves success rate improvements of 53.52\% and 29.41\% for unethical and interactive instructions, respectively. Through extensive experiments and analysis, we highlight the gap between existing mobile agents and real-world user expectations.

移动代理评估基准指令理解AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。