用大模型+新数据集自动从文本生成带并行结构的BPMN流程图
Leveraging Machine Learning and Enhanced Parallelism Detection for BPMN Model Generation from Text
- 结合大语言模型与新标注数据,提升并行流程识别能力
- 在新增15份文档、32个并行网关的数据上训练,准确率显著提升
- 适合需要快速生成规范流程图的企业与自动化系统开发者
高效规划、资源管理和一致运营常依赖将文本流程文档转换为正式的业务流程建模与标注(BPMN)模型。然而,该转换过程仍耗时且成本高昂。现有方法,无论基于规则或机器学习,仍难以应对不同写作风格,且常无法识别流程描述中的并行结构。本文提出一种自动化流程,从文本提取BPMN模型,利用机器学习与大语言模型。关键贡献是构建了一个新标注数据集,对PET数据集扩充了15份新文档,包含32个并行网关,显著改善模型训练效果。该补充使模型更精准捕捉流程中常见的复杂并行结构。所提方法在重建准确性方面表现良好,为组织加速BPMN模型创建提供了有力基础。
原文摘要 · Abstract (English)
Efficient planning, resource management, and consistent operations often rely on converting textual process documents into formal Business Process Model and Notation (BPMN) models. However, this conversion process remains time-intensive and costly. Existing approaches, whether rule-based or machine-learning-based, still struggle with writing styles and often fail to identify parallel structures in process descriptions. This paper introduces an automated pipeline for extracting BPMN models from text, leveraging the use of machine learning and large language models. A key contribution of this work is the introduction of a newly annotated dataset, which significantly enhances the training process. Specifically, we augment the PET dataset with 15 newly annotated documents containing 32 parallel gateways for model training, a critical feature often overlooked in existing datasets. This addition enables models to better capture parallel structures, a common but complex aspect of process descriptions. The proposed approach demonstrates adequate performance in terms of reconstruction accuracy, offering a promising foundation for organizations to accelerate BPMN model creation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。