用多任务学习从病历中自动提取月经特征,准确率超90%。
Multi-Task Learning for Extracting Menstrual Characteristics from Clinical Notes
- 用提示学习结合检索预处理,聚焦病历关键段落。
- 在不足100条标注数据上实现平均90%的F1分数。
- 适合做女性健康研究或临床文本挖掘的开发者参考。
月经健康是女性医疗中的关键但常被忽视的方面。尽管其临床意义重大,结构化医疗记录中却很少包含详细的月经特征数据。为弥补这一空白,我们提出一种新型自然语言处理流程,用于提取月经周期的关键属性:痛经、规律性、经血量和经间期出血。该方法采用GatorTron模型,结合多任务提示学习,并引入混合检索预处理步骤以定位相关文本片段。即使仅在少于100条标注临床笔记上训练,仍优于基线方法,所有月经特征的平均F1得分达到90%。检索步骤显著提升各类方法性能,使模型专注于长篇病历中的相关段落。结果表明,多任务学习与检索结合能有效提升泛化能力与表现,推动临床笔记中月经特征的自动化提取,支持女性健康研究。
原文摘要 · Abstract (English)
Menstrual health is a critical yet often overlooked aspect of women's healthcare. Despite its clinical relevance, detailed data on menstrual characteristics is rarely available in structured medical records. To address this gap, we propose a novel Natural Language Processing pipeline to extract key menstrual cycle attributes -- dysmenorrhea, regularity, flow volume, and intermenstrual bleeding. Our approach utilizes the GatorTron model with Multi-Task Prompt-based Learning, enhanced by a hybrid retrieval preprocessing step to identify relevant text segments. It out- performs baseline methods, achieving an average F1-score of 90% across all menstrual characteristics, despite being trained on fewer than 100 annotated clinical notes. The retrieval step consistently improves performance across all approaches, allowing the model to focus on the most relevant segments of lengthy clinical notes. These results show that combining multi-task learning with retrieval improves generalization and performance across menstrual charac- teristics, advancing automated extraction from clinical notes and supporting women's health research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。