arXiv:2607.05985cs.AIcs.AR2026-07

评估大模型生成设计矩阵的能力,发现其在清晰输入下表现好但易受歧义影响。

Auto-DSM Under the Lens: A Black-Box Evaluation Framework for LLM-Based DSM Generation

论文配图:Auto-DSM Under the Lens: A Black-Box Evaluation Framework for LLM-Based DSM Generation
图 1 · 摘自论文原文
  • 通过黑盒框架对比自动生成与人工验证的设计矩阵
  • 在真实冰箱系统中实现高可重复性,但对模糊表述敏感
  • 为系统工程中的自动化设计提供可审计的评估标准

本文提出一种黑盒评估框架,系统评估大语言模型(LLMs)从结构化技术文档生成设计结构矩阵(DSM)的能力。针对当前自动DSM流程闭源的问题,该框架引入可复现的方法,将生成的DSM(GEN-DSMs)与人工验证的真值矩阵(GT-DSMs)进行对比。评估融合单次运行与多次运行视角,结合结构指标(完整性、正确性、耦合密度)、分类指标(选择性准确率、弃权覆盖率)及稳定性度量(熵、Fleiss' κ)。为此提出综合质量评分(Q)。在虚构抽象系统和真实冰箱分解两个数据集上开展控制实验,涵盖表述差异、参数-数据集对齐及系统复杂度变化。结果表明,LLMs能在结构清晰输入下生成结构合理的DSM并实现高可重复性,但对歧义、依赖定义不一致和提示设计敏感。研究揭示了幻觉与弃权失败的系统性来源,展示了LLM驱动DSM自动化的潜力与局限。该框架为审计Auto-DSM流程提供了透明基准,并为将基于LLM的分解方法集成到模型驱动系统工程(MBSE)工作流中奠定基础。

原文摘要 · Abstract (English)

This paper presents a black-box evaluation framework to systematically assess the ability of Large Language Models (LLMs) to generate Design Structure Matrices (DSMs) from structured technical documentation. Motivated by the closed-source nature of current Auto-DSM pipelines, the framework introduces a reproducible methodology that benchmarks generated DSMs (GEN-DSMs) against manually validated ground-truth matrices (GT-DSMs). The evaluation integrates both single-run and multi-run perspectives, combining structural metrics (Completeness, Correctness, Coupling Density), classification metrics (Selective Accuracy, Abstention Coverage), and stability measures (Entropy, Fleiss' $κ$). To synthesize these aspects, a Composite Quality Score (Q) is proposed. Controlled experiments are conducted on two datasets: a fictive abstract system and a real-world refrigerator decomposition, covering variations in phrasing, parameter-dataset alignment, and system complexity. Results show that LLMs can produce structurally plausible DSMs and achieve high reproducibility under well-structured inputs, but remain sensitive to ambiguity, inconsistent dependency definitions, and prompt formulation. The findings highlight systematic sources of hallucination and abstention failure, demonstrating both the potential and current limitations of LLM-driven DSM automation. The proposed framework provides a transparent benchmark for auditing Auto-DSM pipelines and establishes foundations for integrating LLM-based decomposition methods into model-based systems engineering (MBSE) workflows.

大模型评估设计矩阵系统工程黑盒测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。