arXiv:2604.28181cs.AIcs.CL2026-04被引 1

构建可扩展的虚拟电脑环境,模拟真实长期生产力工作流程。

Synthetic Computers at Scale for Long-Horizon Productivity Simulation

论文配图:Synthetic Computers at Scale for Long-Horizon Productivity Simulation
图 1 · 摘自论文原文
  • 生成带真实文件结构和内容的合成计算机环境。
  • 单次模拟需超8小时运行,平均超过2000步,完成月级任务。
  • 适用于训练具备长程规划能力的智能体,适合研究者与开发者参考。

真实长期生产力工作高度依赖用户特定的计算机环境,其中大量工作上下文通过目录结构和内容丰富的文档、表格、演示文稿等存储与组织。为规模化生成此类生产力场景的合成数据,我们提出「可扩展合成计算机」方法,能够创建具备真实文件层级和丰富内容的虚拟计算机环境。基于每个合成计算机,运行长期模拟:一个智能体设定特定于用户的工作目标,需完成多项专业交付物,相当于约一个月的人工工作量;另一个智能体则扮演该用户,在计算机环境中持续操作——如文件系统导航、与模拟协作方沟通、生成专业文档——直至目标完成。初步实验中,我们构建了1,000个合成计算机,并在其中运行长期模拟,每次模拟需超过8小时的智能体运行时间,平均跨越2,000多步。这些模拟产生丰富的经验学习信号,显著提升智能体在领域内与跨领域生产力评估中的表现。鉴于人物角色可达到百亿规模,该方法在充足算力支持下可扩展至百万甚至千亿级别的合成用户世界,覆盖多样化职业、角色、上下文与需求。我们认为,可扩展的合成计算机构建与大规模模拟,是实现智能体自我进化与代理强化学习在长期生产力任务中发展的关键基础。

原文摘要 · Abstract (English)

Realistic long-horizon productivity work is strongly conditioned on user-specific computer environments, where much of the work context is stored and organized through directory structures and content-rich artifacts. To scale synthetic data creation for such productivity scenarios, we introduce Synthetic Computers at Scale, a scalable methodology for creating such environments with realistic folder hierarchies and content-rich artifacts (e.g., documents, spreadsheets, and presentations). Conditioned on each synthetic computer, we run long-horizon simulations: one agent creates productivity objectives that are specific to the computer's user and require multiple professional deliverables and about a month of human work; another agent then acts as that user and keeps working across the computer -- for example, navigating the filesystem for grounding, coordinating with simulated collaborators, and producing professional artifacts -- until these objectives are completed. In preliminary experiments, we create 1,000 synthetic computers and run long-horizon simulations on them; each run requires over 8 hours of agent runtime and spans more than 2,000 turns on average. These simulations produce rich experiential learning signals, whose effectiveness is validated by significant improvements in agent performance on both in-domain and out-of-domain productivity evaluations. Given that personas are abundant at billion scale, this methodology can in principle scale to millions or even billions of synthetic user worlds with sufficient compute, enabling broader coverage of diverse professions, roles, contexts, environments, and productivity needs. We argue that scalable synthetic computer creation, together with at-scale simulations, is highly promising as a foundational substrate for agent self-improvement and agentic reinforcement learning in long-horizon productivity scenarios.

智能体合成数据长程规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。