探索AI生成完整代码库的挑战与机遇,发现当前工具功能不完善、用户体验差。
Exploring the Challenges and Opportunities of AI-assisted Codebase Generation
- 通过用户实验和访谈,分析开发者使用代码库AI助手的实际交互模式。
- 77%的不满源于功能不满足需求,42%因代码质量差,25%因沟通问题。
- 揭示6大核心挑战与5大应用障碍,为未来设计提供具体方向。
近期的AI代码助手在处理复杂上下文和基于文本描述生成完整代码库方面显著进步,超越了传统的片段级生成。这些代码库AI助手(CBAs)还能扩展或调整代码库,使开发者能聚焦于高层设计与部署决策。尽管已有大量研究关注片段级代码生成,但代码库生成模型仍相对未被充分探索。虽然初期有使用者表示兴奋,但其实际采用率仍低于片段级助手。为更好利用CBAs,需理解开发者如何与之交互,以及它们为何无法满足需求。本文通过一项平衡设计的用户研究与访谈(n=16名学生与开发者),在编码任务中使用CBAs,发现参与者在提示中变化信息:问题描述(48%)、所需功能(98%)、代码结构(48%)及提示撰写过程。尽管策略多样,整体满意度低(均值=2.8,中位数=3,1-5分制)。功能不符是最常见不满原因(77%),其次为代码质量差(42%)与沟通问题(25%)。我们深入剖析不满原因,识别出六个底层挑战,并提炼出五项阻碍集成到工作流中的障碍。最后,调查21个商业CBAs,对比其能力与用户挑战,提出提升效率与实用性的设计机会。
原文摘要 · Abstract (English)
Recent AI code assistants have significantly improved their ability to process more complex contexts and generate entire codebases based on a textual description, compared to the popular snippet-level generation. These codebase AI assistants (CBAs) can also extend or adapt codebases, allowing users to focus on higher-level design and deployment decisions. While prior work has extensively studied the impact of snippet-level code generation, this new class of codebase generation models is relatively unexplored. Despite initial anecdotal reports of excitement about these agents, they remain less frequently adopted compared to snippet-level code assistants. To utilize CBAs better, we need to understand how developers interact with CBAs, and how and why CBAs fall short of developers' needs. In this paper, we explored these gaps through a counterbalanced user study and interview with (n = 16) students and developers working on coding tasks with CBAs. We found that participants varied the information in their prompts, like problem description (48% of prompts), required functionality (98% of prompts), code structure (48% of prompts), and their prompt writing process. Despite various strategies, the overall satisfaction score with generated codebases remained low (mean = 2.8, median = 3, on a scale of one to five). Participants mentioned functionality as the most common factor for dissatisfaction (77% of instances), alongside poor code quality (42% of instances) and communication issues (25% of instances). We delve deeper into participants' dissatisfaction to identify six underlying challenges that participants faced when using CBAs, and extracted five barriers to incorporating CBAs into their workflows. Finally, we surveyed 21 commercial CBAs to compare their capabilities with participant challenges and present design opportunities for more efficient and useful CBAs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。