arXiv:2510.21902cs.SEcs.AI2025-10

测试代码智能体在复杂控制任务中的表现,发现代码访问与动态探索至关重要。

Software Engineering Agents for Embodied Controller Generation : A Study in Minigrid Environments

  • 将代码智能体适配到迷你电网环境,解决20类具身控制任务。
  • 有源码访问时智能体成功率提升40%,动态探索显著优于静态分析。
  • 适合研究具身智能、自动化编程与智能体推理的开发者参考。

软件工程智能体(SWE-Agents)在传统软件任务中表现优异,但在需要高效信息发现的具身任务中表现尚不明确。本文首次对SWE-Agents在具身控制生成任务上的表现进行扩展评估,将Mini-SWE-Agent(MSWEA)适配至迷你电网(Minigrid)环境,解决20种多样化的具身任务。实验对比了不同信息获取条件下的性能表现:是否拥有环境源码访问权限,以及交互式探索能力的差异。结果量化了信息获取水平对智能体性能的影响,并分析了静态代码分析与动态探索在任务求解中的相对重要性。本工作确立具身控制生成为评估SWE-Agents的关键领域,并为高效推理系统研究提供了基准结果。

原文摘要 · Abstract (English)

Software Engineering Agents (SWE-Agents) have proven effective for traditional software engineering tasks with accessible codebases, but their performance for embodied tasks requiring well-designed information discovery remains unexplored. We present the first extended evaluation of SWE-Agents on controller generation for embodied tasks, adapting Mini-SWE-Agent (MSWEA) to solve 20 diverse embodied tasks from the Minigrid environment. Our experiments compare agent performance across different information access conditions: with and without environment source code access, and with varying capabilities for interactive exploration. We quantify how different information access levels affect SWE-Agent performance for embodied tasks and analyze the relative importance of static code analysis versus dynamic exploration for task solving. This work establishes controller generation for embodied tasks as a crucial evaluation domain for SWE-Agents and provides baseline results for future research in efficient reasoning systems.

具身智能代码生成智能体自动化编程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。