arXiv:2506.00150cs.SEcs.AI2025-06

用大模型辅助评估软件架构质量场景,提升决策效率。

Supporting architecture evaluation for ATAM scenarios with LLMs

  • 用大模型分析学生提出的质量场景,自动识别风险点
  • 大模型在多数情况下比人类更准确判断场景风险与权衡
  • 适合架构师、教学者快速筛选关键质量场景

软件架构评估方法长期用于分析设计中不同质量属性之间的权衡。当多个质量属性相互竞争时,难以确定哪些质量场景最应优先处理,也难以对利益相关方提出的需求进行排序。当前评估主要依赖人工,常需长时间头脑风暴来决定最佳场景。为减少工作量并提高评估效率,本文探索使用大语言模型(LLM)部分自动化评估流程。以 MS Copilot 为工具,对比分析学生在软件架构课程中提出的质量场景与大模型的评估结果。初步研究显示,大模型在大多数情况下能生成更优、更准确的风险识别、敏感点分析及权衡判断。这表明生成式 AI 有潜力部分自动化架构评估任务,增强人类决策过程。

原文摘要 · Abstract (English)

Architecture evaluation methods have long been used to evaluate software designs. Several evaluation methods have been proposed and used to analyze tradeoffs between different quality attributes. Having competing qualities leads to conflicts for selecting which quality-attribute scenarios are the most suitable ones that an architecture should tackle and for prioritizing the scenarios required by the stakeholders. In this context, architecture evaluation is carried out manually, often involving long brainstorming sessions to decide which are the most adequate quality scenarios. To reduce this effort and make the assessment and selection of scenarios more efficient, we suggest the usage of LLMs to partially automate evaluation activities. As a first step to validate this hypothesis, this work studies MS Copilot as an LLM tool to analyze quality scenarios suggested by students in a software architecture course and compares the students' results with the assessment provided by the LLM. Our initial study reveals that the LLM produces in most cases better and more accurate results regarding the risks, sensitivity points and tradeoff analysis of the quality scenarios. Overall, the use of generative AI has the potential to partially automate and support the architecture evaluation tasks, improving the human decision-making process.

架构评估大模型应用质量属性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。