arXiv:2510.26969cs.CLcs.AI2025-10中稿 · the LREC 2026 in t…被引 1

用语义框架识别医疗记录中被漏报的性别暴力事件

Frame Semantic Patterns for Identifying Underreporting of Notifiable Events in Healthcare: The Case of Gender-Based Violence

  • 构建8个语义模式,从2100万条葡萄牙语医疗文本中搜寻线索
  • 识别准确率达0.726,有效发现被遗漏的暴力病例
  • 方法透明可复用,适合公共卫生系统中的可解释性NLP应用

我们提出一种在医疗领域识别应报告事件的方法,利用语义框架定义细粒度模式,并在非结构化数据(如电子病历中的开放式文本字段)中进行搜索。该方法应用于初级保健机构中性别暴力(GBV)病例的漏报问题,基于巴西电子SUS APS系统的2100万句葡萄牙语文本语料库,定义并检索8种语义模式。结果经语言学家人工评估,各模式精度均被测量。研究发现,该方法能以0.726的精度有效识别暴力报告,证实其稳健性。该方法设计为透明、高效、低碳且语言无关,可轻松适配其他健康监测场景,推动公共健康系统中NLP技术的伦理与可解释应用。

原文摘要 · Abstract (English)

We introduce a methodology for the identification of notifiable events in the domain of healthcare. The methodology harnesses semantic frames to define fine-grained patterns and search them in unstructured data, namely, open-text fields in e-medical records. We apply the methodology to the problem of underreporting of gender-based violence (GBV) in e-medical records produced during patients' visits to primary care units. A total of eight patterns are defined and searched on a corpus of 21 million sentences in Brazilian Portuguese extracted from e-SUS APS. The results are manually evaluated by linguists and the precision of each pattern measured. Our findings reveal that the methodology effectively identifies reports of violence with a precision of 0.726, confirming its robustness. Designed as a transparent, efficient, low-carbon, and language-agnostic pipeline, the approach can be easily adapted to other health surveillance contexts, contributing to the broader, ethical, and explainable use of NLP in public health systems.

医疗NLP性别暴力语义框架可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。