arXiv:2607.12625cs.CLcs.CV2026-07

让个人助手越用越懂你,跨平台操作更准更快

KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill

论文配图:KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill
图 1 · 摘自论文原文
  • 用记忆和经验分解任务,实现长期规划与精准执行
  • 跨平台无缝切换,长任务成功率达64.1%领先业界
  • 支持多模型通用,知识技能可迁移提升8.5%

OpenClaw作为复杂任务自动化主流框架,存在跨平台GUI交互支持不足及自进化机制缺失问题,限制其在多元设备生态中的适应性与持续优化能力。为此,我们提出‘知深行准’范式,主张用户交互与任务经验积累能直接提升执行准确率与效率,统一认知理解与操作执行。基于此,提出KnowAct-GUIClaw,采用知-路-行-反思框架,解决OpenClaw的GUI操作短板与跨平台、递归自进化瓶颈。主机代理利用累积经验与任务知识进行长周期任务分解与分配(知);可插拔的GUI子代理配备经验驱动的记忆系统(知)与自演化技能库(行),支持跨平台迁移与快速集成。框架持续存储用户画像与反馈,提升任务分解与工具调用准确率。在Android、iOS、HarmonyOS和Windows上的实验表明,该框架显著提升效率、准确率与跨平台适应性。尤其,使用开源Kimi-2.6模型的GUIClaw在长周期MobileWorld基准上达到64.1%最高性能,超越所有同类框架及闭源模型如Seed-2.0-Pro与GPT-5.5。此外,框架支持的知识记忆与执行技能可在不同基础模型间迁移,使Kimi-2.6性能提升8.5%。

原文摘要 · Abstract (English)

OpenClaw has emerged as a leading agent framework for complex task automation, yet it faces insufficient cross-platform GUI interaction support and a well-built self-evolution mechanism. These flaws limit its adaptation to diverse device ecosystems and prevent performance improvements through continuous learning from execution experience. To resolve these issues, we propose the Know Deeply, Act Perfectly paradigm for personal assistants, which holds that accumulated user interaction and task-running experience directly improve execution accuracy and efficiency, unifying cognitive comprehension and operational execution. Based on this paradigm, we introduce KnowAct-GUIClaw, a novel Know-Route-Act-Reflect framework designed to address OpenClaw's GUI manipulation deficits and break through its cross-platform and recursive self-improvement constraints. First, the host agent leverages accumulated interaction experience and task-relevant knowledge for long-horizon task decomposition and allocation (Know). Second, a pluggable GUI subagent with an experience-attributable memory system (Know) and self-evolving skill library (Act), enabling seamless cross-platform migration and fast-path integration. Especially, this framework continuously stores user profiles and feedback to improve the accuracy of task decomposition and tool calls. Extensive experiments across Android, iOS, HarmonyOS and Windows show that KnowAct-GUIClaw achieves superior efficiency, accuracy and cross-platform adaptability. Especially, the GUIClaw with open-source Kimi-2.6 models achieves the best performance (64.1%) on the long-horizon MobileWorld benchmark, beating all agentical frameworks and closed-source agentical models, e.g., Seed-2.0-Pro and GPT-5.5. Additionally, the knowledgeable memory and execution skills supported by our framework are transferable across diverse base models, improving by 8.5% with Kimi-2.6.

个人助理GUI自动化自进化跨平台

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。