arXiv:2503.09896cs.CLcs.AI2025-03被引 4

基于规则的共指消解系统在医学文本中表现优异,准确率达89.6%。

A Rule Based Solution to Co-reference Resolution in Clinical Text

  • 通过人工观察训练数据构建规则,实现医学文本共指链识别
  • 在多个医学数据集上达到89.6%的整体性能
  • 适合需要高可解释性的临床语义分析场景

目标:本研究旨在构建一个针对生物医学领域的有效共指消解系统。材料与方法:实验所用数据来自2011年i2b2自然语言处理挑战赛,该挑战涉及医疗文档中的共指消解任务。临床文本中的概念提及已被标注,需在每篇文档中将具有共指关系的提及链接成共指链。通常有两种构建自动共指链接系统的方法:一是手工构建规则,二是使用机器学习模型从训练数据中自动学习,并在测试数据上执行消解任务。结果:实验表明,现有共指消解系统能够发现部分共指链接,而本研究提出的基于规则的系统在多数共指链接识别中表现良好。系统在多个医学数据集上取得89.6%的整体性能。结论:实验结果表明,基于对训练数据的观察手动设计规则,是实现该生物医学领域共指消解任务高性能的有效方法。

原文摘要 · Abstract (English)

Objective: The aim of this study was to build an effective co-reference resolution system tailored for the biomedical domain. Materials and Methods: Experiment materials used in this study is provided by the 2011 i2b2 Natural Language Processing Challenge. The 2011 i2b2 challenge involves coreference resolution in medical documents. Concept mentions have been annotated in clinical texts, and the mentions that co-refer in each document are to be linked by coreference chains. Normally, there are two ways of constructing a system to automatically discover co-referent links. One is to manually build rules for co-reference resolution, and the other category of approaches is to use machine learning systems to learn automatically from training datasets and then perform the resolution task on testing datasets. Results: Experiments show the existing co-reference resolution systems are able to find some of the co-referent links, and our rule based system performs well finding the majority of the co-referent links. Our system achieved 89.6% overall performance on multiple medical datasets. Conclusion: The experiment results show that manually crafted rules based on observation of training data is a valid way to accomplish high performance in this coreference resolution task for the critical biomedical domain.

共指消解医学文本规则系统自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。