arXiv:2604.07717cs.CLcs.AI2026-04

用大模型从临床记录中自动识别艾滋病污名,助力改善患者照护。

Detecting HIV-Related Stigma in Clinical Narratives Using Large Language Models

  • 基于大语言模型构建文本分析工具,识别四种艾滋病相关污名类型。
  • GatorTron-large在零样本下表现最佳,微平均F1达0.62,少样本提示显著提升生成模型效果。
  • 工具可辅助医疗人员发现患者心理困扰,适合临床研究与公共卫生干预使用。

艾滋病相关污名是影响艾滋病感染者(PLWH)心理健康、就医参与度及治疗结果的重要心理社会因素。尽管污名经历存在于临床笔记中,但缺乏现成的提取与分类工具。本研究旨在开发基于大语言模型(LLM)的工具,从临床笔记中识别艾滋病污名。数据来自佛罗里达大学健康中心2012至2022年间接受治疗的感染者临床记录。通过专家标注的污名关键词并结合临床词嵌入迭代扩展,共筛选出1,332条候选句子,并由专家人工标注四个污名子量表:公众态度担忧、披露顾虑、负面自我形象和个性化污名。比较了GatorTron-large与BERT等编码型基线模型,以及GPT-OSS-20B、LLaMA-8B和MedGemma-27B等生成型大模型在零样本与少样本提示下的表现。结果显示,GatorTron-large整体表现最优(微平均F1 = 0.62)。少样本提示显著提升生成模型性能,5样本提示下GPT-OSS-20B与LLaMA-8B的微平均F1分别为0.57和0.59。不同污名子量表表现差异明显,其中负面自我形象最易预测,个性化污名仍最困难。零样本生成推理存在显著失败率(最高达32%)。本研究首次构建了可用于临床实践的艾滋病污名自然语言处理工具。

原文摘要 · Abstract (English)

Human immunodeficiency virus (HIV)-related stigma is a critical psychosocial determinant of health for people living with HIV (PLWH), influencing mental health, engagement in care, and treatment outcomes. Although stigma-related experiences are documented in clinical narratives, there is a lack of off-the-shelf tools to extract and categorize them. This study aims to develop a large language model (LLM)-based tool for identifying HIV stigma from clinical notes. We identified clinical notes from PLWH receiving care at the University of Florida (UF) Health between 2012 and 2022. Candidate sentences were identified using expert-curated stigma-related keywords and iteratively expanded via clinical word embeddings. A total of 1,332 sentences were manually annotated across four stigma subscales: Concern with Public Attitudes, Disclosure Concerns, Negative Self-Image, and Personalized Stigma. We compared GatorTron-large and BERT as encoder-based baselines, and GPT-OSS-20B, LLaMA-8B, and MedGemma-27B as generative LLMs, under zero-shot and few-shot prompting. GatorTron-large achieved the best overall performance (Micro F1 = 0.62). Few-shot prompting substantially improved generative model performance, with 5-shot GPT-OSS-20B and LLaMA-8B achieving Micro-F1 scores of 0.57 and 0.59, respectively. Performance varied by stigma subscale, with Negative Self-Image showing the highest predictability and Personalized Stigma remaining the most challenging. Zero-shot generative inference exhibited non-trivial failure rates (up to 32%). This study develops the first practical NLP tool for identifying HIV stigma in clinical notes.

自然语言处理艾滋病临床文本分析大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。