arXiv:2501.12538cs.CLcs.AI2025-01

用AI分析7000+新冠后遗症病例报告,发现数据严重缺少数十种社会因素。

Academic case reports lack diversity: Assessing the presence and diversity of sociodemographic and behavioral factors related to Post COVID-19 Condition

  • 通过NLP技术从病例报告中提取26类社会健康因素
  • 发现年龄、医疗可及性等常见,但种族、住房状况等被严重忽略
  • 适合关注健康不平等与医疗数据偏见的研究者

了解脆弱群体中新冠后遗症(PCC)的流行率、差异及症状变化对改善照护和解决交叉不平等问题至关重要。本研究旨在通过自然语言处理(NLP)技术构建整合社会决定健康因素(SDOH)的框架,分析PCC病例报告中各类因素的代表性与多样性。基于超过7,000份来自LitCOVID数据库的病例报告,构建了PCC病例报告语料库,并对其中709份报告进行标注,涵盖26类核心SDOH实体类型。采用预训练命名实体识别(NER)模型、人工审核与数据增强提升质量。开发集成NER、自然语言推理(NLI)、三元组与频率分析的NLP流水线。评估了编码器单向变压器模型与RNN模型在NER任务上的表现,微调后的BERT模型在应对不同句式结构和类别稀疏性方面优于传统RNN。探索性分析显示实体丰富度存在差异,常见实体包括病情、年龄和医疗可及性,而种族与住房状况等敏感类别严重不足。三元组分析揭示年龄、性别与病情常共现。NLI分析表明,如‘曾遭受暴力或虐待’和‘有医疗保险’的蕴含率高达82.4%至80.3%,而‘女性身份’‘已婚’‘患终末期疾病’等属性则呈现高矛盾率(70.8%至98.5%)。

原文摘要 · Abstract (English)

Understanding the prevalence, disparities, and symptom variations of Post COVID-19 Condition (PCC) for vulnerable populations is crucial to improving care and addressing intersecting inequities. This study aims to develop a comprehensive framework for integrating social determinants of health (SDOH) into PCC research by leveraging NLP techniques to analyze disparities and variations in SDOH representation within PCC case reports. Following construction of a PCC Case Report Corpus, comprising over 7,000 case reports from the LitCOVID repository, a subset of 709 reports were annotated with 26 core SDOH-related entity types using pre-trained named entity recognition (NER) models, human review, and data augmentation to improve quality, diversity and representation of entity types. An NLP pipeline integrating NER, natural language inference (NLI), trigram and frequency analyses was developed to extract and analyze these entities. Both encoder-only transformer models and RNN-based models were assessed for the NER objective. Fine-tuned encoder-only BERT models outperformed traditional RNN-based models in generalizability to distinct sentence structures and greater class sparsity. Exploratory analysis revealed variability in entity richness, with prevalent entities like condition, age, and access to care, and underrepresentation of sensitive categories like race and housing status. Trigram analysis highlighted frequent co-occurrences among entities, including age, gender, and condition. The NLI objective (entailment and contradiction analysis) showed attributes like "Experienced violence or abuse" and "Has medical insurance" had high entailment rates (82.4%-80.3%), while attributes such as "Is female-identifying," "Is married," and "Has a terminal condition" exhibited high contradiction rates (70.8%-98.5%).

健康不平等NLP应用数据偏见新冠后遗症

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。