arXiv:2604.20523cs.SEcs.AI2026-04被引 1

用大模型直接分析半形式化产品线蓝图,实现早期验证。

Early-Stage Product Line Validation Using LLMs: A Study on Semi-Formal Blueprint Analysis

  • 直接在自然语言描述的蓝图上运行特征模型分析
  • 推理优化模型准确率达88%-89%,接近专业求解器
  • 适合快速验证产品线架构,提升开发效率

我们研究大型语言模型(LLMs)能否直接对半形式化的文本蓝图(即特征层次与约束的简洁约束语言描述)执行特征模型分析操作(AOs),以支持软件产品线规划阶段的早期验证。使用12个最先进的LLMs和16种标准分析操作,将结果与基于求解器的基准工具FLAMA进行对比。结果显示,推理优化型模型(如Grok 4 Fast Reasoning、Gemini 2.5 Pro)在所有测试蓝图和操作上平均准确率达88%-89%,接近求解器的正确性。研究识别出结构解析与约束推理中的系统性错误,并揭示了准确率与计算成本之间的权衡,为模型选择提供依据。这些发现表明,LLMs可作为轻量级助手,用于早期变体验证。

原文摘要 · Abstract (English)

We study whether Large Language Models (LLMs) can perform feature model analysis operations (AOs) directly on semi-formal textual blueprints, i.e., concise constrained-language descriptions of feature hierarchies and constraints, enabling early validation in Software Product Line scoping. Using 12 state-of-the-art LLMs and 16 standard AOs, we compare their outputs against the solver-based oracle FLAMA. Results show that reasoning-optimized models (e.g., Grok 4 Fast Reasoning, Gemini 2.5 Pro) achieve 88-89% average accuracy across all evaluated blueprints and operations, approaching solver correctness. We identify systematic errors in structural parsing and constraint reasoning, and highlight accuracy-cost trade-offs that inform model selection. These findings position LLMs as lightweight assistants for early variability validation.

产品线大模型验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。