arXiv:2606.18293cs.SEcs.AI2026-06

评估用自然语言编程能否替代写代码,探索AI生成软件的可行性

Vibe Coding Ate My Homework: An evaluation of AI approaches to greenfield software engineering and programming

论文配图:Vibe Coding Ate My Homework: An evaluation of AI approaches to greenfield software engineering and programming
图 1 · 摘自论文原文
  • 用自然语言指令完成简单独立的Python编程任务
  • 发现当前大模型在复杂逻辑实现上仍存明显缺陷
  • 适合关注AI编程潜力与局限的研究者和开发者

得益于生成式AI的快速发展,我们正经历一场可能永久改变人机交互方式的范式变革。自然语言提示被广泛用于构建应用与编码基础设施,无需领域知识,这种现象被称为「vibe coding」。它被视为编程自诞生以来不断追求更高抽象层次的终点——完全摒弃代码语法,转而使用母语进行编程。本文旨在评估vibe coding在全新项目(greenfield)软件工程任务中的可行性,并分析现有评测基准。为此,我们构建了一个评估套件,用于分析大语言模型在执行简单的、孤立的Python编程任务时的表现,以获得对该问题的有限但明确的洞察。

原文摘要 · Abstract (English)

Thanks to rapid developments in generative AI, we are in the midst of a paradigm shift that may change how we interact with computers forever. We have observed a growth in the use of natural language prompts to build applications and coding infrastructures without underlying knowledge of the field, and this practice has been dubbed `vibe coding.' It arguably represents what the field of programming has been building towards since the beginning, with every higher level of abstraction that is conceived. Vibe coding promises to be the endpoint for the meta of high-level programming as far as method of input is concerned: eliminating a human's use of code syntax entirely in favour of programming in their mother tongue. This paper aims to evaluate the viability of vibe coding for greenfield software engineering tasks, as well as analyse the benchmarks that have been used to measure its software engineering prowess. To this end, we have developed an evaluation suite for analysing an LLM's proficiency in carrying out simple, isolated greenfield programming tasks in Python to provide scoped insight on the matter.

AI编程自然语言编程大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。