arXiv:2605.02244cs.SEcs.AI2026-05

构建三元数据集,训练能处理长期复杂工程任务的智能助手

The Conversations Beneath the Code: Triadic Data for Long-Horizon Software Engineering Agents

  • 采集人类工程师对话、人机协作与跨职能项目协同的同步数据
  • 提出四层评估框架,确保数据质量可验证、可复现
  • 适合研究长周期智能软件工程代理的学者与团队

当前前沿软件工程智能体在短期任务上表现饱和,却难以应对资深工程师才需处理的长期、多角色、需求模糊的任务。本文主张下一代智能体的训练数据不应仅依赖更大规模的GitHub爬取或单一智能体轨迹,而应采用三元数据:同步记录人类工程师间的协作对话、人机交互过程以及持续数周的跨职能工作成果。我们提出两种核心数据产品:基于刺激回忆协议采集的长期专家任务轨迹,以及模拟跨职能公司的虚拟团队(含资深工程师、产品经理、设计师、数据科学家)在共享基础设施上完成模糊需求交付的仿真环境。进一步构建了四层证据体系,包括机械验证、统计特征分析、探针实验和预注册盲评,以保障数据集质量。论证该数据可在12-18个月内通过相邻领域成熟方法获取,并是解决四个关键智能体训练难题的实证钥匙,建议学界将此类数据建设纳入近期研究议程。

原文摘要 · Abstract (English)

Frontier software engineering agents have saturated short-horizon benchmarks while regressing on the work that constitutes senior engineering: long-horizon, multi-engineer, ambiguous-specification deliverables. This paper takes a position on what training data is needed to close the gap. The substrate for the next generation of SWE agents is neither larger GitHub scrapes nor more solo-agent trajectories nor -- sufficient by itself -- open human-AI dialogue logs. It is triadic data: synchronized capture of the human-human conversations where engineering context is formed, the human-AI sessions where that context is partially consumed, and the multi-week cross-functional work that surrounds both. We argue that the canonical instantiation of triadic data is two complementary products: long-horizon expert trajectories captured under stimulated-recall protocols, and simulated cross-functional companies -- instrumented teams of senior engineers, product managers, designers, and data scientists working through ambiguous deliverables on shared infrastructure. We further specify a four-tier evidence framework through which any such corpus -- triadic or otherwise -- must justify its quality to a fine-tuning researcher: mechanical verification, statistical corpus characterization, probe experiments, and pre-registered blind evaluation. We argue that this data is capturable in 12-18 months with methods already mature in adjacent fields, that it is the empirical key to four open questions in agent training, and that the field's near-term research agenda should include it explicitly.

软件工程智能代理三元数据长周期任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。