arXiv:2502.19518cs.SEcs.AI2025-02被引 7

测试大模型对iOS VIPER架构的理解与生成能力,发现其擅长设计创新但难记细节。

Assessing LLMs for Front-end Software Architecture Knowledge

  • 用布鲁姆分类法构建评估框架,覆盖从记忆到创造的六个认知层级
  • 在生成和评估任务中表现优秀,但精确复现架构细节能力不足
  • 为大模型在软件架构领域的应用提供可复用的评测基准,适合开发者参考

大型语言模型(LLMs)在自动化软件开发任务中展现出巨大潜力,但在软件设计任务中的能力仍不明确。本研究考察了大模型在理解、复现和生成复杂 iOS 应用程序设计模式 VIPER 架构方面的能力。我们基于布鲁姆分类法构建了一个全面的评估框架,用于衡量模型在记忆、理解、应用、分析、评价和创造等不同认知层级的表现。实验使用 ChatGPT 4 Turbo 2024-04-09 版本进行,结果表明,该模型在高阶任务如评价和创造上表现优异,但在需要精确检索架构细节的低阶任务中存在困难。这些发现凸显了大模型降低开发成本的潜力,也揭示了其在真实软件设计场景中有效应用的障碍。本研究提出了一种评估大模型在软件架构领域能力的基准格式,旨在推动更稳健、易用的AI驱动开发工具的发展。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated significant promise in automating software development tasks, yet their capabilities with respect to software design tasks remains largely unclear. This study investigates the capabilities of an LLM in understanding, reproducing, and generating structures within the complex VIPER architecture, a design pattern for iOS applications. We leverage Bloom's taxonomy to develop a comprehensive evaluation framework to assess the LLM's performance across different cognitive domains such as remembering, understanding, applying, analyzing, evaluating, and creating. Experimental results, using ChatGPT 4 Turbo 2024-04-09, reveal that the LLM excelled in higher-order tasks like evaluating and creating, but faced challenges with lower-order tasks requiring precise retrieval of architectural details. These findings highlight both the potential of LLMs to reduce development costs and the barriers to their effective application in real-world software design scenarios. This study proposes a benchmark format for assessing LLM capabilities in software architecture, aiming to contribute toward more robust and accessible AI-driven development tools.

大模型软件架构评测基准iOS开发

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。