用智能推荐和可视化降低文档信息抽取标注难度
Assisted Data Annotation for Business Process Information Extraction from Textual Documents
- 通过推荐系统识别文本中的流程信息,辅助人工标注
- 实验显示工作量减少51%,标注质量提升38.9%
- 适合需要高效构建流程数据集的研究者与企业
基于机器学习的自然语言流程描述生成流程模型,可缓解业务流程发现阶段耗时耗力的问题。然而,该研究受限于缺乏大规模高质量数据集,主要因缺乏有效工具支持数据创建,导致工作负荷大、数据质量差。本文探索两种辅助功能:文本中流程信息的推荐系统,以及已识别信息的图形化流程模型可视化。针对31名参与者的控制实验表明,推荐系统可将各项工作量降低最高51.0%,显著提升标注质量最高达38.9%。所有数据与代码均已公开,以推动新型辅助策略研究。
原文摘要 · Abstract (English)
Machine-learning based generation of process models from natural language text process descriptions provides a solution for the time-intensive and expensive process discovery phase. Many organizations have to carry out this phase, before they can utilize business process management and its benefits. Yet, research towards this is severely restrained by an apparent lack of large and high-quality datasets. This lack of data can be attributed to, among other things, an absence of proper tool assistance for dataset creation, resulting in high workloads and inferior data quality. We explore two assistance features to support dataset creation, a recommendation system for identifying process information in the text and visualization of the current state of already identified process information as a graphical business process model. A controlled user study with 31 participants shows that assisting dataset creators with recommendations lowers all aspects of workload, up to $-51.0\%$, and significantly improves annotation quality, up to $+38.9\%$. We make all data and code available to encourage further research on additional novel assistance strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。