用人类操作认知数据训练AI,让其夜间自动完成复杂办公任务。
PC Agent: While You Sleep, AI Works -- A Cognitive Journey into Digital World

- 通过记录用户操作全过程,还原思维过程以训练AI。
- 仅用133条高质量认知数据,就能完成跨50步的多应用复杂任务。
- 适合想开发智能数字助手的研究者与开发者。
设想一个世界:AI可在你睡眠时处理工作,如整理研究资料、撰写报告或制作明日所需演示文稿。然而,当前数字代理虽能执行简单任务,仍难以应对人类日常面临的复杂真实工作。我们提出PC Agent,一种通过人类认知迁移实现该愿景的关键进展。核心洞察在于,从执行简单“任务”到处理复杂“工作”的路径,在于高效捕捉并学习用户使用计算机时的认知过程。为验证此假设,我们引入三项关键创新:(1) PC Tracker,一种轻量级基础设施,可高效收集带有完整认知上下文的高质量人机交互轨迹;(2) 两阶段认知补全流程,将原始交互数据转化为富含语义动作与思维过程的丰富认知轨迹;(3) 多代理系统,结合规划代理进行决策,以及接地代理实现鲁棒的视觉定位。初步实验显示,在幻灯片创建任务中,仅需133条认知轨迹,PC Agent即可处理涉及最多50步、跨多个应用的复杂工作场景。这表明方法具备高度数据效率,凸显训练强大数字代理的关键在于采集人类认知数据。我们已开源完整框架,包括数据采集基础设施与认知补全方法,旨在降低研究社区构建真正智能数字代理的门槛。
原文摘要 · Abstract (English)
Imagine a world where AI can handle your work while you sleep - organizing your research materials, drafting a report, or creating a presentation you need for tomorrow. However, while current digital agents can perform simple tasks, they are far from capable of handling the complex real-world work that humans routinely perform. We present PC Agent, an AI system that demonstrates a crucial step toward this vision through human cognition transfer. Our key insight is that the path from executing simple "tasks" to handling complex "work" lies in efficiently capturing and learning from human cognitive processes during computer use. To validate this hypothesis, we introduce three key innovations: (1) PC Tracker, a lightweight infrastructure that efficiently collects high-quality human-computer interaction trajectories with complete cognitive context; (2) a two-stage cognition completion pipeline that transforms raw interaction data into rich cognitive trajectories by completing action semantics and thought processes; and (3) a multi-agent system combining a planning agent for decision-making with a grounding agent for robust visual grounding. Our preliminary experiments in PowerPoint presentation creation reveal that complex digital work capabilities can be achieved with a small amount of high-quality cognitive data - PC Agent, trained on just 133 cognitive trajectories, can handle sophisticated work scenarios involving up to 50 steps across multiple applications. This demonstrates the data efficiency of our approach, highlighting that the key to training capable digital agents lies in collecting human cognitive data. By open-sourcing our complete framework, including the data collection infrastructure and cognition completion methods, we aim to lower the barriers for the research community to develop truly capable digital agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。