先感知再推理,让手机助手更准更省力地主动帮忙。
Perceive Before Reasoning: A Pre-Reasoning Perception Framework for Efficient and Reliable Proactive Mobile Agents

- 先用轻量感知模块判断是否该干预,再决定是否启动复杂推理。
- 在基准测试中误触发率降低,成功率和推理效率均提升。
- 适合追求高效可靠的主动式移动助手研发者使用。
多模态大语言模型(MLLM)显著推动了移动代理的发展,但主动辅助仍面临挑战:代理需在决定如何协助前判断何时介入。现有系统常将两个决策整合于单一MLLM流程中,导致保守的干预过滤与全面的协助生成之间目标错位,且在应保持沉默时仍产生冗余推理。为此,我们提出预推理感知框架(PRPF),采用分两阶段设计,先感知后推理。PRPF引入轻量级多模态主动感知器(MPP)实现干预门控与上下文压缩,并仅在需要干预时激活主动推理器(PAR)。在ProactiveMobile基准上的实验表明,相比ProactiveMobile基线,PRPF显著降低了误触发率(FTR),同时提升了成功率(SR)与推理效率。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have substantially advanced mobile agents, yet proactive mobile assistance remains challenging because agents must decide when to intervene before determining how to assist. Existing systems often implement these two decisions within a unified MLLM-based pipeline, leading to goal misalignment between conservative intervention filtering and comprehensive assistance generation, as well as redundant inference when the agent should remain silent. To address these limitations, we propose the Pre-Reasoning Perception Framework (PRPF), a two-stage framework built on perceiving before reasoning. PRPF introduces a lightweight Multimodal Proactive Perceptor (MPP) for intervention gating and context compression, and activates the Proactive Agent Reasoner (PAR) only when intervention is warranted. Experiments on the ProactiveMobile benchmark show that PRPF substantially reduces false trigger rates (FTR) while improving success rates (SR) and inference efficiency over the ProactiveMobile baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。