arXiv:2606.19637cs.CLcs.AI2026-06

临床文本中的自杀风险标注受数据构建方式影响,不能直接当真金白银的标签用。

Before the Labels: How Dataset Construction Shapes Suicidality Detection in Clinical Text

  • 用MIMIC-III病历构建标注集时,医生判断、编码规则和单人标注导致标签带有主观性。
  • 相同标签背后可能包含不同时间性、否定或不确定性的临床表达,差异大却混为一谈。
  • 适合关注临床文本标注可信度的研究者和医疗AI开发者阅读。

临床自然语言处理越来越多依赖电子健康记录(EHR)数据检测自杀行为,常将临床文书视为比社交媒体更可靠的“真实标签”。本文认为,这种说法掩盖了基于EHR的自杀风险数据集实际上是对自杀风险的特定操作化定义,其形成受书写者身份、事件边界设定及模糊性处理方式的影响。以ScAN数据集为例,该数据集基于MIMIC-III临床笔记构建,我们发现治理限制、基于ICD的队列筛选、单标注员标注以及住院期级别聚合,共同导致标签反映的是临床医生的判断,将自杀风险视为有边界的事件,并默认意图可从记录中可靠推断。语言学分析显示,相同标签下包含的时间性、否定和不确定性表达存在显著异质性。因此,临床NLP研究应在解读标签前,先审视其背后隐含的假设。

原文摘要 · Abstract (English)

Clinical NLP increasingly relies on electronic health record (EHR) data to detect suicidal behaviors, treating clinical documentation as more reliable ground truth than social media. We argue that this framing obscures how EHR-based suicidality datasets encode a particular operationalization of suicidality, shaped by who authors the data, how episodes are bounded, and how ambiguity is resolved. We ground this argument in a case study of the ScAN dataset, built over MIMIC-III clinical notes. We show how governance constraints, ICD-based cohort selection, single-annotator labeling, and hospital-stay-level aggregation produce labels that reflect clinician-documented judgments, treat suicidality as a bounded episode, and assume that intent can be reliably inferred from documentation. A linguistic analysis demonstrates that identical labels subsume heterogeneous clinical framings differing in temporality, negation, and uncertainty. We argue that clinical NLP should examine the assumptions embedded in suicidality datasets before interpreting their labels as ground truth.

临床NLP自杀检测数据标注医疗文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。