让智能代理主动感知用户深层需求并长期记忆,实时做出精准干预。
PASK: Toward Intent-Aware Proactive Agents with Long-Term Memory
- 通过流式意图识别与多层记忆建模实现主动响应
- 在延迟约束下性能媲美Gemini3-Flash,更深入识别用户意图
- 适用于需长期理解、实时决策的智能助手场景
主动性是通用人工智能的核心期待。现有工作多局限于实验室环境,难以应对真实世界中深度、复杂性、模糊性、精度及实时性等多重挑战。本文研究此类场景:有效干预需从持续上下文中推断隐含需求,并在延迟与长周期约束下,基于动态用户记忆执行动作。我们提出通用范式DD-MM-PAS(需求检测、记忆建模、主动代理系统),并基于此构建Pask系统。其中包含流式意图识别模型IntentFlow,混合记忆结构(工作区、用户、全局记忆),以及代理系统框架,形成闭环。同时,构建了基于用户授权数据、经数千轮人工精修的LatentNeeds-Bench真实世界基准。实验表明,IntentFlow在延迟约束下表现接近顶尖的Gemini3-Flash模型,且能更准确捕捉深层用户意图。
原文摘要 · Abstract (English)
Proactivity is a core expectation for AGI. Prior work remains largely confined to laboratory settings, leaving a clear gap in real-world proactive agent: depth, complexity, ambiguity, precision and real-time constraints. We study this setting, where useful intervention requires inferring latent needs from ongoing context and grounding actions in evolving user memory under latency and long-horizon constraints. We first propose DD-MM-PAS (Demand Detection, Memory Modeling, Proactive Agent System) as a general paradigm for streaming proactive AI agent. We instantiate this paradigm in Pask, with streaming IntentFlow model for DD, a hybrid memory (workspace, user, global) for long-term MM, PAS infra framework and introduce how these components form a closed loop. We also introduce LatentNeeds-Bench, a real-world benchmark built from user-consented data and refined through thousands of rounds of human editing. Experiments show that IntentFlow matches leading Gemini3-Flash models under latency constraints, while identifying deeper user intent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。