通过真实代码库中的探索轨迹,揭示SWE智能体的思考与决策模式。
Projecting the Emerging Mindset of SWE Agent by Launching a Wild Code Understanding Journey

- 构建可记录的代码理解框架Ada,限定工具接口实现可观测探索。
- 408条轨迹显示不同模型在效率、路径多样性与认知根基上的差异。
- 提供观察智能体行为的方法论,适合研究代码生成与推理的学者。
软件工程智能体(SWE agents)在真实代码仓库中通过工具引导的路径工作,但其行为难以用具体可观测的方式描述。这些轨迹记录了工具使用、中间推理、证据选择和自主停止,却无法解释为何做出特定操作、信任哪些证据或何时认为理解足够。这种张力使轨迹数据既有限又宝贵:当通过严谨观察解读时,可成为研究智能体行为的实证基础。我们提出Ada,一个用于仓库级代码理解的受限框架。Ada通过有限工具接口进入真实代码库,允许开放式探索,同时保持轨迹可记录性。在这一开放但受限的环境中,Ada决定查看位置、精读内容、整合部分理解及关闭仓库评估。通过导航、证据选择、综合、定位与停止等观察视角,我们使智能体的思维-行动链可见,避免简化为工具调用次数或猜测隐藏意图。结合多个模型、仓库、任务类型与启动条件,408条轨迹揭示了智能体在效率、路径多样性、认知根基与干预边界上的差异,为真实代码库中观测SWE智能体行为提供了方法学基础。
原文摘要 · Abstract (English)
Software engineering agents (SWE agents) increasingly work through tool-mediated trajectories in real repositories, yet their behavior remains difficult to characterize in concrete, observable terms. These trajectories record tool use, intermediate reasoning, evidence selection, and self-directed stopping, but they do not by themselves explain why particular moves were chosen, what evidence was trusted, or when understanding was judged sufficient. This tension makes trajectory data both limited and valuable: faithful, replayable traces can become an empirical substrate for studying agent behavior when interpreted through disciplined observation. We introduce Ada, a scoped apparatus for repository-level code understanding. Ada enters real codebases through a bounded tool interface, allowing open-ended exploration to remain recordable as finite trajectories. Across this wild-but-bounded setting, Ada chooses where to look, what to read closely, when to consolidate partial understanding, and when to close its account of the repository. We project Ada's think-action chains through observation lenses that make navigation, evidence selection, synthesis, grounding, and stopping visible without reducing behavior to raw tool counts or speculating about hidden intent. Read together, these lenses produce behavioral profiles grounded in recorded movement through software worlds. Across 408 trajectories, spanning multiple models, repositories, task families, and launch conditions, the study shows how faithful digital traces can be transformed into disciplined, comparable projections of emerging SWE-agent mindset. The results expose differences in efficiency, trajectory diversity, epistemic grounding, and the limits of intervention, while providing a methodological foundation for observing SWE agent behavior in real codebases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。