arXiv:2603.08013cs.AI2026-03被引 5

让AI从被动执行变为主动预测用户操作,提升智能助手的前瞻性能力。

PIRA-Bench: A Transition from Reactive GUI Agents to GUI-based Proactive Intent Recommendation Agents

  • 基于连续视觉输入构建多意图交错的复杂轨迹数据集
  • 在真实噪声和多任务切换场景下实现高精度意图推荐
  • 适合研究主动式人机交互与多模态大模型的应用者

当前图形用户界面(GUI)代理主要采用被动响应范式:用户必须提供明确指令才能执行任务。而一个真正的智能AI助手应具备主动性,能直接从连续视觉输入(如移动端或桌面截图)中感知用户意图,并在无需显式提示时提供及时建议。这一范式转变面临巨大挑战——真实屏幕活动极少线性,常包含长时程轨迹、噪音浏览、无意义操作及多任务切换。为此,我们提出PIRA-Bench(主动意图推荐代理基准),用于评估多模态大语言模型(MLLMs)在连续、弱监督视觉输入下的表现。与传统被动数据集不同,PIRA-Bench包含多个交错意图和带有不同用户画像上下文的噪声片段,要求代理在识别可行动事件的同时契合用户偏好。此外,我们提出PIRF基线框架,一种具备记忆机制与状态追踪能力的模型,使通用MLLM能够管理多任务线程并应对误导性视觉输入。PIRA-Bench是迈向鲁棒、主动式基于GUI个人助手的重要一步。

原文摘要 · Abstract (English)

Current Graphical User Interface (GUI) agents operate primarily under a reactive paradigm: a user must provide an explicit instruction for the agent to execute a task. However, an intelligent AI assistant should be proactive, which is capable of anticipating user intentions directly from continuous visual inputs, such as mobile or desktop screenshots, and offering timely recommendations without explicit user prompting. Transitioning to this proactive paradigm presents significant challenges. Real-world screen activity is rarely linear; it consists of long-horizon trajectories fraught with noisy browsing, meaningless actions, and multithreaded task-switching. To address this gap, we introduce PIRA-Bench (Proactive Intent Recommendation Agent Benchmark), a novel benchmark for evaluating multimodal large language models (MLLMs) on continuous, weakly-supervised visual inputs. Unlike reactive datasets, PIRA-Bench features complex trajectories with multiple interleaved intents and noisy segments with various user profile contexts, challenging agents to detect actionable events while fitting to user preferences. Furthermore, we propose the PIRF baseline, a memory-aware, state-tracking framework that empowers general MLLMs to manage multiple task threads and handle misleading visual inputs. PIRA-Bench serves as an initial step toward robust and proactive GUI-based personal assistants.

主动推荐GUI智能多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。