用规则+机器学习混合策略,让信息抽取更高效可复现。
Development of the user-friendly decision aid Rule-based Evaluation and Support Tool (REST) for optimizing the resources of an information extraction task
- 先人工标注少量数据,再用工具自动判断每类实体该用规则还是机器学习。
- 在12类实体上验证,工具能准确预测规则开发难度和提取效果。
- 适合希望减少标注工作、提升系统可解释性的信息抽取团队。
与机器学习(ML)和大语言模型(LLM)相比,规则在可持续性、可迁移性、可解释性和开发负担方面具有优势。本文提出一种可持续的规则与机器学习结合的信息抽取(IE)方法。流程始于专家在单次会话中对代表性数据子集进行手动标注。我们开发并验证了名为规则评估与支持工具(REST)的决策辅助工具,帮助标注员为每个实体任务决定是否采用规则作为默认方案或使用机器学习。REST使标注员能够可视化每类实体在自由文本中的特征、规则构建可行性及预期性能指标。机器学习仅作为备选方案,从而大幅减少训练所需的标注工作量。在12类实体的应用场景中,REST表现出良好的可复现性,证明其外部有效性。
原文摘要 · Abstract (English)
Rules could be an information extraction (IE) default option, compared to ML and LLMs in terms of sustainability, transferability, interpretability, and development burden. We suggest a sustainable and combined use of rules and ML as an IE method. Our approach starts with an exhaustive expert manual highlighting in a single working session of a representative subset of the data corpus. We developed and validated the feasibility and the performance metrics of the REST decision tool to help the annotator choose between rules as a by default option and ML for each entity of an IE task. REST makes the annotator visualize the characteristics of each entity formalization in the free texts and the expected rule development feasibility and IE performance metrics. ML is considered as a backup IE option and manual annotation for training is therefore minimized. The external validity of REST on a 12-entity use case showed good reproducibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。