构建首个移动端主动智能评估基准,推动手机代理从被动执行到主动预测进化。
ProactiveMobile: A Comprehensive Benchmark for Boosting Proactive Intelligence on Mobile Devices
- 通过多维度上下文信号推断用户隐含意图,生成可执行的API操作序列。
- 覆盖3660+实例、14种真实场景,支持多答案标注与专家质量审核。
- 实测表明当前大模型主动能力普遍不足,但可通过训练提升,适合研究者参考。
多模态大语言模型在移动代理开发中取得显著进展,但其能力仍主要局限于被动响应模式,仅执行用户明确指令。主动智能新范式要求代理自主预判需求并主动发起行动,是移动代理的未来方向。然而,该领域发展受限于缺乏能应对真实复杂性且支持客观可执行评估的基准。为此,我们提出ProactiveMobile,一个系统性推进该领域研究的综合性基准。ProactiveMobile将主动任务形式化为:基于设备端四维上下文信号推断隐含用户意图,并从包含63个API的全面函数池中生成可执行函数序列。基准涵盖超过3,660个实例、14种现实场景,采用多答案标注以体现真实复杂性。为确保质量,由30名专家团队进行最终审计,验证事实准确性、逻辑一致性和动作可行性,并修正不合规条目。大量实验表明,微调后的Qwen2.5-VL-7B-Instruct成功率达19.15%,优于o1(15.71%)和GPT-5(7.39%)。结果表明主动能力是当前多模态大模型普遍缺失的关键能力,但具有可学习性,凸显本基准在主动智能评估中的重要价值。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have made significant progress in mobile agent development, yet their capabilities are predominantly confined to a reactive paradigm, where they merely execute explicit user commands. The emerging paradigm of proactive intelligence, where agents autonomously anticipate needs and initiate actions, represents the next frontier for mobile agents. However, its development is critically bottlenecked by the lack of benchmarks that can address real-world complexity and enable objective, executable evaluation. To overcome these challenges, we introduce ProactiveMobile, a comprehensive benchmark designed to systematically advance research in this domain. ProactiveMobile formalizes the proactive task as inferring latent user intent across four dimensions of on-device contextual signals and generating an executable function sequence from a comprehensive function pool of 63 APIs. The benchmark features over 3,660 instances of 14 scenarios that embrace real-world complexity through multi-answer annotations. To ensure quality, a team of 30 experts conducts a final audit of the benchmark, verifying factual accuracy, logical consistency, and action feasibility, and correcting any non-compliant entries. Extensive experiments demonstrate that our fine-tuned Qwen2.5-VL-7B-Instruct achieves a success rate of 19.15%, outperforming o1 (15.71%) and GPT-5 (7.39%). This result indicates that proactivity is a critical competency widely lacking in current MLLMs, yet it is learnable, emphasizing the importance of the proposed benchmark for proactivity evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。