arXiv:2609.04013cs.AIcs.LG2026-09

用大模型零样本筛查肾病,数据少也能用。

LLM4CKD: Large Language Models for Early Stage Chronic Kidney Disease Screening

论文配图:LLM4CKD: Large Language Models for Early Stage Chronic Kidney Disease Screening
图 1 · 摘自论文原文
  • 不用训练,用结构化提示词直接推理
  • 少量样本下表现媲美甚至超越传统模型
  • 适合数据稀缺场景,但稳定性依赖模型

早期筛查慢性肾病(CKD)对及时干预至关重要,但多数机器学习(ML)和深度学习(DL)方法需标注数据与模型训练,限制了其在真实筛查中的应用。本研究评估了大语言模型(LLMs)在零样本和少样本上下文学习下的CKD筛查效果,并与传统ML、DL方法及表格基础模型(TFM)进行对比。我们提出一个框架,采用临床筛选的表格特征与结构化提示模板,实现无需任务特定训练的LLM推理。在多种提示风格、特征配置和数据设置下评估LLM性能,结果表明:在少量样本条件下,LLMs可达到竞争性表现,常匹配或优于传统方法;但其性能受模型影响较大,输入复杂度升高时稳定性下降。相比之下,ML、DL和TFM模型随训练数据增加表现出更一致的提升。总体而言,研究揭示了数据效率与稳定性之间的权衡,表明在标注数据有限时,LLMs可作为灵活的补充筛查工具。

原文摘要 · Abstract (English)

Early screening of chronic kidney disease (CKD) is critical for timely intervention, yet most machine learning (ML) and deep learning (DL) approaches require labeled data and model training, limiting their use in real-world screening settings. This study evaluates the effectiveness of large language models (LLMs) for CKD screening under zero-shot and few-shot in-context learning settings and compares them with traditional ML and DL methods. We propose a framework that uses clinically selected tabular features and structured prompt templates to enable LLM-based inference without task-specific training. LLM performance is evaluated across multiple prompt styles, feature configurations, and data settings, and compared with standard ML, DL, and tabular foundation model (TFM) baselines, and existing CKD screening tools. The results show that LLMs can achieve competitive performance using only a small number of examples, often matching or outperforming traditional approaches in low-data settings. However, their performance remains model-dependent and less stable as input complexity increases. In contrast, ML, DL, and TFM models show more consistent improvement with larger training data. Overall, the findings highlight a trade-off between data efficiency and stability, suggesting that LLMs may serve as a flexible complementary approach for CKD screening when labeled data are limited.

大模型医疗筛查少样本学习肾病

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。