用大模型从患者自述中提取健康社会决定因素,提升糖尿病风险预测。
Structured Insight from Unstructured Data: Large Language Models for SDOH-Driven Diabetes Risk Prediction
- 用大模型分析65名老年患者的未结构化访谈,生成可量化的社会健康数据
- 基于提取特征与传统指标联合建模,实现对糖尿病控制水平的60%准确预测
- 为临床提供可扩展的非结构化数据转化方案,适合医疗决策支持系统
健康的社会决定因素(SDOH)在2型糖尿病(T2D)管理中至关重要,但常缺失于电子病历和风险预测模型中。现有个体级SDOH数据多通过结构化筛查工具收集,难以捕捉患者经历的复杂性及诊所人群的独特需求。本研究探索使用大语言模型(LLMs)从65名年龄≥65岁的T2D患者未结构化访谈中提取结构化SDOH信息,并评估提取特征及原始叙事文本对糖尿病控制状态的预测价值。采用检索增强生成技术,将访谈内容转化为简洁、可操作的定性摘要与结构化定量评分,用于风险建模。这些评分独立及与传统实验室生物标志物结合,输入岭回归、套索、随机森林和XGBoost等机器学习模型。同时,评估多个LLM直接从红移A1C值的访谈文本中预测患者糖尿病控制水平(低/中/高)的能力,结果达到60%准确率。研究表明,LLMs能有效将非结构化SDOH数据转化为结构化洞察,为临床风险模型和决策提供可扩展的新路径。
原文摘要 · Abstract (English)
Social determinants of health (SDOH) play a critical role in Type 2 Diabetes (T2D) management but are often absent from electronic health records and risk prediction models. Most individual-level SDOH data is collected through structured screening tools, which lack the flexibility to capture the complexity of patient experiences and unique needs of a clinic's population. This study explores the use of large language models (LLMs) to extract structured SDOH information from unstructured patient life stories and evaluate the predictive value of both the extracted features and the narratives themselves for assessing diabetes control. We collected unstructured interviews from 65 T2D patients aged 65 and older, focused on their lived experiences, social context, and diabetes management. These narratives were analyzed using LLMs with retrieval-augmented generation to produce concise, actionable qualitative summaries for clinical interpretation and structured quantitative SDOH ratings for risk prediction modeling. The structured SDOH ratings were used independently and in combination with traditional laboratory biomarkers as inputs to linear and tree-based machine learning models (Ridge, Lasso, Random Forest, and XGBoost) to demonstrate how unstructured narrative data can be applied in conventional risk prediction workflows. Finally, we evaluated several LLMs on their ability to predict a patient's level of diabetes control (low, medium, high) directly from interview text with A1C values redacted. LLMs achieved 60% accuracy in predicting diabetes control levels from interview text. This work demonstrates how LLMs can translate unstructured SDOH-related data into structured insights, offering a scalable approach to augment clinical risk models and decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。