arXiv:2507.08969cs.CL2025-07

用AI分析病历文本,发现特定人群病历中歧视性语言更多

Application of CARE-SD text classifier tools to assess distribution of stigmatizing and doubt-marking language features in EHR

  • 通过词典匹配和分类器识别病历中的质疑性与歧视性表述
  • 黑人患者、低收入群体病历中歧视语言使用率高出1.16至2.46倍
  • 护士和社会工作者记录中歧视性语言更突出,适合医疗公平研究者阅读

电子健康记录(EHR)是医疗团队中患者污名化传播的关键媒介。本研究通过扩展词典匹配与监督学习分类器,在MIMIC-III数据集中识别出怀疑标记和歧视性标签的语言特征。利用泊松回归模型评估各语言特征的预测因子。结果显示,非裔患者(相对风险RR: 1.16)、拥有医疗保险或政府保险的患者(RR: 2.46)、自费患者(RR: 2.12)以及患有多种被污名化疾病或精神健康问题的患者,其病历中歧视性标签出现频率更高;怀疑性语言模式类似,男性患者使用更多(RR: 1.25)。护士(RR: 1.40)和社工(RR: 2.25)使用的歧视性语言也显著偏高。结论表明,历史上受污名化的患者群体面临更高频率的歧视性语言,且由多类医务人员共同加剧。

原文摘要 · Abstract (English)

Introduction: Electronic health records (EHR) are a critical medium through which patient stigmatization is perpetuated among healthcare teams. Methods: We identified linguistic features of doubt markers and stigmatizing labels in MIMIC-III EHR via expanded lexicon matching and supervised learning classifiers. Predictors of rates of linguistic features were assessed using Poisson regression models. Results: We found higher rates of stigmatizing labels per chart among patients who were Black or African American (RR: 1.16), patients with Medicare/Medicaid or government-run insurance (RR: 2.46), self-pay (RR: 2.12), and patients with a variety of stigmatizing disease and mental health conditions. Patterns among doubt markers were similar, though male patients had higher rates of doubt markers (RR: 1.25). We found increased stigmatizing labels used by nurses (RR: 1.40), and social workers (RR: 2.25), with similar patterns of doubt markers. Discussion: Stigmatizing language occurred at higher rates among historically stigmatized patients, perpetuated by multiple provider types.

医疗公平病历分析语言偏见自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。