arXiv:2608.04148cs.SEcs.AI2026-08

新手通过扮演四种角色,在多智能体修复系统中学习与智能体协作。

AgentForge: An Immersive Role-Playing Platform for Learning Agentic Software Engineering

论文配图:AgentForge: An Immersive Role-Playing Platform for Learning Agentic Software Engineering
图 1 · 摘自论文原文
  • 设计沉浸式角色扮演平台,分四角色模拟真实开发流程。
  • 37名新手参与,完成率高,代码评审任务最耗时且最困难。
  • 提升对软件修复和人机协作的理解,适合初学者快速上手。

代理型人工智能在软件开发中日益用于协调规划、实现、审查和测试,但其决策与交互过程透明度有限。许多系统假设用户能有效引导AI并验证输出,这对新手构成挑战——他们需同时学习代理工作原理、如何有效协作及批判性评估输出。为此,我们提出 extit{AgentForge},一个沉浸式学习系统:新手扮演四种软件工程角色之一(任务规划者、补丁作者、代码审查者或测试执行者),在多智能体代码修复流程中实践。每个练习中,新手执行所选角色,其余三角色由AI代理完成。通过基于角色的支架支持与元认知辅助,AgentForge明确各角色职责,使协作过程与中间产物可视化,并鼓励新手监控与评估自身决策。对37名新手开发者的研究表明,借助AI代理支持,参与者任务完成率高;但不同角色交互需求差异显著:代码审查任务需更多交互轮次、路径重定向和更长时间(p_{\mathrm{adj}} = .004),被视作最具挑战性。尽管如此,参与者在软件修复与代理协作理解方面均有显著提升(p_{\mathrm{adj}} < .001)。结果表明,AgentForge可帮助新手在实践中发展软件工程能力,并更批判、高效地与代理型AI协作。

原文摘要 · Abstract (English)

Agentic AI is increasingly used to coordinate planning, implementation, review, and testing in software development, yet it often offers limited transparency into its decisions and interactions. Many such systems also assume that users can effectively guide the AI's decisions and validate its outputs. This assumption poses a particular challenge for novices, who must simultaneously learn how agentic AI works, how to collaborate with it effectively, and how to evaluate its outputs critically. To address this challenge, we present \textit{AgentForge}, an immersive learning system in which novices take on one of four software-engineering roles: Task Planner, Patch Author, Code Reviewer, or Test Runner, within a multi-agent code-repair workflow. In each practice session, the novices perform their chosen role while AI agents perform the remaining three. Through role-based scaffolding and metacognitive support, AgentForge clarifies role-specific responsibilities, makes agent coordination and intermediate artifacts visible, and encourages novices to monitor and evaluate their decisions. In a study with 37 novice developers, participants achieved high task-completion rates with AI-agent support. However, interaction demands differed significantly across practices: the Code Reviewer practice required more interaction turns, reroutes, and completion time ($p_{\mathrm{adj}} = .004$) and was perceived as the most challenging. Participants nevertheless reported significant gains in their understanding of software repair and agent collaboration ($p_{\mathrm{adj}} < .001$). These findings suggest that AgentForge can help novices develop practical software-engineering skills while learning to collaborate with agentic AI more critically and effectively.

智能体编程教育人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。