评测大模型生成流程图的能力,发现其语法和实用性强但语义理解仍有不足。
Assessing the Business Process Modeling Competences of Large Language Models
- 构建四维评估框架,从语法、实用、语义和有效性四个角度衡量流程图质量。
- 大模型在语法和实用性上接近人类专家,但在语义理解和正确性上仍落后。
- 为大模型优化提供方向,适合关注AI辅助流程设计的从业者参考。
业务流程建模与表示(BPMN)的创建是一项复杂且耗时的任务,需要领域知识和对建模范式的熟练掌握。近年来,大型语言模型(LLMs)在直接从自然语言生成BPMN模型方面取得进展,超越了早期的文本到流程方法,在处理复杂描述方面表现出更强能力。然而,目前缺乏对LLM生成流程图的系统性评估。现有工作或采用“大模型作为评判者”的方式,或未考虑模型质量的既定维度。为此,我们提出BEF4LLM——一个包含语法质量、实用质量、语义质量和有效性的四视角评估框架。利用该框架,我们对开源大模型进行了全面分析,并将其性能与人类建模专家进行对比。结果显示,大模型在语法和实用质量方面表现优异,而人类在语义方面更具优势;但两者得分差距相对较小,凸显了大模型在流程建模中的竞争潜力,尽管在有效性和语义质量方面仍存在挑战。研究揭示了当前大模型在流程建模中的优势与局限,为未来模型优化和微调提供了指导,对推动大模型在实际业务流程建模中的应用至关重要。
原文摘要 · Abstract (English)
The creation of Business Process Model and Notation (BPMN) models is a complex and time-consuming task requiring both domain knowledge and proficiency in modeling conventions. Recent advances in large language models (LLMs) have significantly expanded the possibilities for generating BPMN models directly from natural language, building upon earlier text-to-process methods with enhanced capabilities in handling complex descriptions. However, there is a lack of systematic evaluations of LLM-generated process models. Current efforts either use LLM-as-a-judge approaches or do not consider established dimensions of model quality. To this end, we introduce BEF4LLM, a novel LLM evaluation framework comprising four perspectives: syntactic quality, pragmatic quality, semantic quality, and validity. Using BEF4LLM, we conduct a comprehensive analysis of open-source LLMs and benchmark their performance against human modeling experts. Results indicate that LLMs excel in syntactic and pragmatic quality, while humans outperform LLMs in semantic aspects; however, the differences in scores are relatively modest, highlighting LLMs' competitive potential despite challenges in validity and semantic quality. The insights highlight current strengths and limitations of using LLMs for BPMN modeling and guide future model development and fine-tuning. Addressing these areas is essential for advancing the practical deployment of LLMs in business process modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。