arXiv:2511.14967cs.SEcs.AI2025-11被引 5

首个评估自然语言生成序列图的基准,专为代码工程设计。

MermaidSeqBench: An Evaluation Benchmark for NL-to-Mermaid Sequence Diagram Generation

  • 混合人工验证与模型增强构建132个测试样本
  • 用LLM作裁判,多维度评估语法、异常处理等正确性
  • 适合研究大模型结构化生成能力的学者和工程师

大型语言模型(LLMs)在从自然语言描述生成结构化图表方面展现出巨大潜力,尤其在软件工程中生成Mermaid序列图。然而,缺乏针对该任务的评测基准,阻碍了对模型能力的严谨、系统性评估与比较。为此,我们提出MermaidSeqBench,一个经人工验证并结合合成扩展的基准,用于评估LLM从自然语言提示生成Mermaid序列图的能力。该基准包含132个样本,通过人工验证流程、LLM增强和规则扩展的混合方法构建。评估采用LLM-as-a-judge模型,从语法正确性、激活处理、异常处理及实际可用性等多个细粒度维度进行评价。为验证基准的有效性与灵活性,我们在多个主流LLM上进行了初步评估,并使用多种LLM裁判,揭示了不同模型与评估模式间的显著能力差距。MermaidSeqBench为结构化图表生成提供了评估基础,推动了对大模型在结构化生成任务中能力与局限性的科学理解。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown great promise in generating structured diagrams from natural language descriptions, particularly Mermaid sequence diagrams for software engineering. However, the lack of existing benchmarks to assess the LLM's correctness on this task hinders rigorous, systematic evaluation and principled comparison of model capabilities on this task. To address this shortcoming, we introduce MermaidSeqBench, a human-verified and synthetically extended benchmark for assessing LLM capabilities in generating Mermaid sequence diagrams from natural language prompts. The benchmark consists of 132 samples developed via a hybrid methodology of human-verified flows, LLM-based augmentation, and rule-based expansion. The evaluation uses an LLM-as-a-judge model to assess generation across various fine-grained metrics such as syntax correctness, activation handling, error handling, and practical usability. To demonstrate the effectiveness and flexibility of our benchmark, we perform initial evaluations on numerous state-of-the-art LLMs with multiple LLM judges which reveal significant capability gaps across models and evaluation modes. MermaidSeqBench provides a foundation for evaluating structured diagram generation and advances the scientific understanding of LLM capabilities and limitations in structured generation tasks.

序列图自然语言生成大模型评估软件工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。