构建细粒度医学术语简化任务,助力公众理解复杂医疗文本。
JEBS: A Fine-grained Biomedical Lexical Simplification Task
- 将医学术语简化拆解为识别、分类、生成三步,实现精准处理。
- 发布包含21,595条替换的JEBS数据集,覆盖400篇生物医学摘要。
- 适合作为医学信息可读性研究与自然语言处理应用的基准工具。
在线医学文献使健康信息前所未有的易获取,但复杂的医学术语仍阻碍普通公众理解。尽管已有生物医学文本简化用的平行和可比语料库,但它们混淆了简化过程中的多种句法与词汇操作。为支持更精准的开发与评估,我们提出细粒度词汇简化任务及数据集——医学简化术语解释(JEBS,https://github.com/bill-from-ri/JEBS-data)。JEBS任务包括识别复杂术语、分类替换方式、生成替代文本三个子任务。数据集包含400篇生物医学摘要及其人工简化版本中,共10,314个术语的21,595条替换。同时提供基于规则与Transformer模型的三类子任务基线结果。JEBS任务、数据及基线结果为医学术语替换或解释系统的开发与严格评估奠定了基础。
原文摘要 · Abstract (English)
Online medical literature has made health information more available than ever, however, the barrier of complex medical jargon prevents the general public from understanding it. Though parallel and comparable corpora for Biomedical Text Simplification have been introduced, these conflate the many syntactic and lexical operations involved in simplification. To enable more targeted development and evaluation, we present a fine-grained lexical simplification task and dataset, Jargon Explanations for Biomedical Simplification (JEBS, https://github.com/bill-from-ri/JEBS-data ). The JEBS task involves identifying complex terms, classifying how to replace them, and generating replacement text. The JEBS dataset contains 21,595 replacements for 10,314 terms across 400 biomedical abstracts and their manually simplified versions. Additionally, we provide baseline results for a variety of rule-based and transformer-based systems for the three sub-tasks. The JEBS task, data, and baseline results pave the way for development and rigorous evaluation of systems for replacing or explaining complex biomedical terms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。