用少量易得的临床数据,让大模型零样本筛查慢性肾病。
From Many to Meaningful: Feature-Guided Zero-Shot Chronic Kidney Disease Screening Using Large Language Models

- 选关键临床特征,用提示模板转成文本输入大模型。
- 在三个国家数据集上,准确率显著提升且结果稳定。
- 适合资源有限地区,无需训练即可部署使用。
早期筛查慢性肾病(CKD)对防止不可逆进展至关重要;然而,许多机器学习(ML)筛查方法因依赖大量标注数据、高成本病理检测或高维临床特征,难以在社区和资源匮乏环境中部署,且对人群和分布变化的鲁棒性差。本研究探讨了在零样本设置下使用大语言模型(LLMs)进行早期CKD筛查的可行性。我们提出一种特征引导的零样本框架,仅使用一组临床有意义且易于获取的社区级特征,而非全部临床变量。特征选择通过机器学习分析确定,形成精简且临床相关的变量子集。将表格患者记录通过标准化提示模板序列化为文本,实现零样本推理。评估了四种LLM(LLaMA-3、Qwen-3、Mistral、GPT-4o-mini)在完整特征集与选定子集上的表现。在覆盖三个国家的三组异质性CKD数据集上评估泛化能力。所有模型与数据集组合中,选定特征集均带来一致且统计显著的平衡准确率与概率估计提升,达到筛查可用水平。结果表明,大模型可基于最少社区可及患者特征,实现无需训练的临床有意义筛查,为真实世界筛查场景提供实用补充。
原文摘要 · Abstract (English)
Early screening of chronic kidney disease (CKD) is essential for preventing irreversible progression; however, many machine learning (ML)-based screening methods remain difficult to deploy in community and resource-limited screening settings due to their reliance on large labeled datasets, resource-intensive pathology tests, or high-dimensional clinical features, and limited robustness to population and distributional shifts. This study examines the feasibility of using large language models (LLMs) for early-stage CKD screening in a zero-shot setting, without dataset-specific training. We propose a feature-guided zero-shot framework that evaluates LLM performance using a selected set of clinically meaningful, readily available community-based features, rather than exhaustive clinical inputs. Feature selection was guided by ML-based analysis to identify a compact, clinically relevant subset of variables. Tabular patient records were subsequently serialized into text using standardized prompt templates to enable zero-shot inference. The zero-shot performance of four LLMs (LLaMA-3, Qwen-3, Mistral, and GPT-4o-mini) was evaluated using both the full feature set and the selected subset. Generalizability was assessed across three heterogeneous CKD datasets spanning three countries. Across models and datasets, the selected feature set yielded consistent and statistically significant improvements in balanced accuracy and probability estimates, achieving performance levels suitable for screening purposes. These findings suggest that LLMs can support clinically meaningful, training-free CKD screening using minimal community-accessible patient features, offering a practical complement to conventional ML methods in real-world screening contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。