arXiv:2505.10831cs.HCcs.AI2025-05被引 73

用电脑行为生成通用用户模型,让助手主动理解并预判需求。

Creating General User Models from Computer Use

  • 通过观察屏幕截图等多模态数据,自动推断用户意图和偏好。
  • 能准确识别婚礼筹备、协作受阻等复杂情境,置信度可量化。
  • 适合做智能助手、主动提醒、跨应用自适应的交互系统。

人机交互长期追求技术理解用户——从偏好习惯到日常行为的目的与时机。然而现有用户模型仍碎片化、局限于特定应用,难以实现灵活推理。本文提出通用用户模型(GUM),通过观察用户任意计算机交互(如设备截图)来学习其知识与偏好。GUM将非结构化观察转化为带置信度的命题,可推断用户正在为参加婚礼准备,或因合作者反馈而多次停滞编辑。该架构支持从多模态输入中生成新命题,检索上下文命题,并持续更新已有认知。我们展示了其在增强聊天助手上下文、智能管理通知、跨应用适配交互代理等方面的应用。还构建了主动助手(GUMBO),利用GUM自主发现并执行用户潜在需求。评估显示,GUM能做出校准且准确的推断,基于其的助手可主动识别并完成用户未明确提出的操作。总体而言,GUM利用多模态模型理解非结构化上下文,推动人机交互愿景落地,开启新型预测性交互系统。

原文摘要 · Abstract (English)

Human-computer interaction has long imagined technology that understands us-from our preferences and habits, to the timing and purpose of our everyday actions. Yet current user models remain fragmented, narrowly tailored to specific apps, and incapable of the flexible reasoning required to fulfill these visions. This paper presents an architecture for a general user model (GUM) that learns about you by observing any interaction you have with your computer. The GUM takes as input any unstructured observation of a user (e.g., device screenshots) and constructs confidence-weighted propositions that capture user knowledge and preferences. GUMs can infer that a user is preparing for a wedding they're attending from messages with a friend. Or recognize that a user is struggling with a collaborator's feedback on a draft by observing multiple stalled edits and a switch to reading related work. GUMs introduce an architecture that infers new propositions about a user from multimodal observations, retrieves related propositions for context, and continuously revises existing propositions. To illustrate the breadth of applications that GUMs enable, we demonstrate how they augment chat-based assistants with context, manage OS notifications to selectively surface important information, and enable interactive agents that adapt to preferences across apps. We also instantiate proactive assistants (GUMBOs) that discover and execute useful suggestions on a user's behalf using their GUM. In our evaluations, we find that GUMs make calibrated and accurate inferences about users, and that assistants built on GUMs proactively identify and perform actions that users wouldn't think to request explicitly. Altogether, GUMs introduce methods that leverage multimodal models to understand unstructured context, enabling long-standing visions of HCI and entirely new interactive systems that anticipate user needs.

用户建模主动助手多模态理解人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。