评估大模型辅助流程建模工具的人类中心体验,发现可用性尚可但信任度低。
Human-Centered Evaluation of an LLM-Based Process Modeling Copilot: A Mixed-Methods Study with Domain Experts
- 结合专家访谈与问卷,用混合方法评估大模型流程助手的使用体验。
- 用户可用性评分67.2分(满分100),信任度仅48.8分,可靠性最受关注。
- 适合流程建模专家、企业质量保障团队,需提升追问深度以改善输出质量。
将大语言模型(LLMs)融入业务流程管理工具有望使非专家也能掌握业务流程建模与标注(BPMN)。尽管自动化框架可评估语法和语义质量,却忽略了信任、易用性和专业契合度等人类因素。我们通过焦点小组和标准化问卷,对五位流程建模专家开展了混合方法评估,检验了自研的基于LLM的BPMN协作者。结果揭示出可用性感知尚可(均值CUQ得分67.2/100),但信任度显著偏低(均值48.8/100),其中可靠性为最核心关切(均值1.8/5)。此外,还发现输出质量问题、使用障碍,并强调需让大模型提出更深入的澄清性问题。我们提出了从领域专家支持到企业质量保证在内的五种应用场景。研究证明,人类中心评估必须补充自动化基准测试,以全面评估大模型建模代理的有效性。
原文摘要 · Abstract (English)
Integrating Large Language Models (LLMs) into business process management tools promises to democratize Business Process Model and Notation (BPMN) modeling for non-experts. While automated frameworks assess syntactic and semantic quality, they miss human factors like trust, usability, and professional alignment. We conducted a mixed-methods evaluation of our proposed solution, an LLM-powered BPMN copilot, with five process modeling experts using focus groups and standardized questionnaires. Our findings reveal a critical tension between acceptable perceived usability (mean CUQ score: 67.2/100) and notably lower trust (mean score: 48.8\%), with reliability rated as the most critical concern (M=1.8/5). Furthermore, we identified output-quality issues, prompting difficulties, and a need for the LLM to ask more in-depth clarifying questions about the process. We envision five use cases ranging from domain-expert support to enterprise quality assurance. We demonstrate the necessity of human-centered evaluation complementing automated benchmarking for LLM modeling agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。