arXiv:2601.07779cs.MAcs.AI2026-01ACL被引 19

让电脑操作智能体更稳健通用,解决长流程任务中记忆丢失和新场景适应难题。

OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agent

  • 用里程碑式长期记忆实现任务级自我修正,防止视觉信息丢失。
  • 通过实时网页搜索生成视觉对齐教程,提升陌生场景下的操作精度。
  • 在多个基准上达到新最好性能,尤其在OSWorld上达65.84%成功率。

尽管视觉语言模型(VLM)推动了计算机使用智能体(CUAs)的发展,现有框架在长周期工作流中的鲁棒性以及新领域泛化能力仍显不足。其根源在于缺乏对历史视觉上下文的细粒度控制,以及缺乏视觉感知的教程检索机制。为此,我们提出OS-Symphony,一个整体性框架,由协调器驱动两大创新:(1) 反思-记忆智能体,采用里程碑驱动的长期记忆机制,实现轨迹级自我修正,有效缓解长周期任务中的视觉上下文丢失问题;(2) 多功能工具智能体,包含采用SeeAct范式的多模态搜索引擎,可在基于浏览器的沙箱环境中导航,动态合成与当前视觉状态一致的实时教程,从而解决未知场景下的保真度问题。实验结果表明,OS-Symphony在不同模型规模下均带来显著性能提升,在三个在线基准上取得新最佳表现,尤其在OSWorld上达到65.84%。

原文摘要 · Abstract (English)

While Vision-Language Models (VLMs) have significantly advanced Computer-Using Agents (CUAs), current frameworks struggle with robustness in long-horizon workflows and generalization in novel domains. These limitations stem from a lack of granular control over historical visual context curation and the absence of visual-aware tutorial retrieval. To bridge these gaps, we introduce OS-Symphony, a holistic framework that comprises an Orchestrator coordinating two key innovations for robust automation: (1) a Reflection-Memory Agent that utilizes milestone-driven long-term memory to enable trajectory-level self-correction, effectively mitigating visual context loss in long-horizon tasks; (2) Versatile Tool Agents featuring a Multimodal Searcher that adopts a SeeAct paradigm to navigate a browser-based sandbox to synthesize live, visually aligned tutorials, thereby resolving fidelity issues in unseen scenarios. Experimental results demonstrate that OS-Symphony delivers substantial performance gains across varying model scales, establishing new state-of-the-art results on three online benchmarks, notably achieving 65.84% on OSWorld.

智能体长程任务视觉推理自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。