用未来提问法测试大模型能否当隐式医学世界模型。
Future Querying: Can LLMs Serve as Implicit Medical World Models?

- 让大模型回答患者未来临床问题,无需特定任务训练。
- 小模型本地微调后效果接近大公司闭源系统。
- 适合注重隐私、需本地部署的医疗场景。
传统临床预测模型依赖特定任务流程和结构化数据,难以扩展且未充分利用非结构化文本。为此,我们提出未来提问法,评估大语言模型(LLMs)作为隐式医学世界模型的能力,即回答关于患者未来时间点的临床问题。该框架基于非结构化临床记录,采用端点无关训练,使单一模型能在不进行人工特征工程或任务重训的情况下,回答患者病程中的多样化临床问题。在新构建的合成医学报告数据集及真实MIMIC-IV重症监护数据上的实验表明,小型本地微调的开源模型性能可媲美甚至接近大型专有系统,证明该框架适用于注重隐私、可本地部署的场景。结果表明,大模型能捕捉临床动态的部分特征。
原文摘要 · Abstract (English)
Traditional clinical prediction models rely on task-specific pipelines and curated, structured data, which scale poorly and underutilize unstructured text. To address this, we introduce future querying, a paradigm that probes whether large language models (LLMs) can function as implicit medical world models by evaluating their ability to answer time-indexed clinical queries about a patient's future. Our framework operates on unstructured clinical documentation using endpoint-agnostic training, enabling a single model to answer diverse clinical queries over patient trajectories without manual feature engineering or task-specific retraining. We show that small, locally fine-tuned open-weight models can match or approach larger proprietary systems, making the framework suitable for privacy-preserving, on-premise deployment. Evaluated on a new synthetic medical reports dataset and real ICU notes from the MIMIC-IV dataset, our results provide encouraging evidence that LLMs can capture aspects of clinical dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。