Surfer 2用视觉感知实现跨平台通用操作,性能远超此前系统。
Surfer 2: The Next Generation of Cross-Platform Computer Use Agents
- 基于视觉输入的统一架构,无需环境专用接口
- 在四大基准上均达领先准确率,多试次超人类表现
- 适合需要跨平台自动化的人工智能研究与应用
构建能泛化至网页、桌面和移动端的智能体仍是开放挑战,因以往系统依赖环境特定接口,限制了跨平台部署。我们提出Surfer 2,一种纯视觉观测驱动的统一架构,在三类环境中均达到当前最优性能。该系统融合分层上下文管理、解耦规划与执行,以及自验证与自适应恢复机制,支持长任务周期下的可靠运行。在WebVoyager上达97.1%准确率,WebArena上69.6%,OSWorld上60.1%,AndroidWorld上87.1%,显著优于所有先前方法且无需任务微调。多次尝试下,其性能超越人类水平。结果表明,系统性调度可放大基础模型能力,仅通过视觉交互即可实现通用计算机控制,同时呼吁下一代视觉语言模型以实现成本效益的帕累托最优。
原文摘要 · Abstract (English)
Building agents that generalize across web, desktop, and mobile environments remains an open challenge, as prior systems rely on environment-specific interfaces that limit cross-platform deployment. We introduce Surfer 2, a unified architecture operating purely from visual observations that achieves state-of-the-art performance across all three environments. Surfer 2 integrates hierarchical context management, decoupled planning and execution, and self-verification with adaptive recovery, enabling reliable operation over long task horizons. Our system achieves 97.1% accuracy on WebVoyager, 69.6% on WebArena, 60.1% on OSWorld, and 87.1% on AndroidWorld, outperforming all prior systems without task-specific fine-tuning. With multiple attempts, Surfer 2 exceeds human performance on all benchmarks. These results demonstrate that systematic orchestration amplifies foundation model capabilities and enables general-purpose computer control through visual interaction alone, while calling for a next-generation vision language model to achieve Pareto-optimal cost-efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。