提出CalTwin方法,提升医疗世界模型在数据分布变化下的可靠性与预测可信度。
CalTwin: Towards Calibrated, Shift-Robust Medical World Models via Fisher-Information Regularisation

- 通过费雪信息正则化缓解数据分布偏移问题,增强模型泛化能力。
- 显著降低外部测试集上9.1%的未来状态预测误差,提升鲁棒性。
- 适合临床部署中需高可信度与跨机构适用性的医疗生成模型研究者。
医疗世界模型旨在学习患者或器官生理的潜在状态及其随干预变化的演化规律,支持从影像诊断到数字孪生治疗规划等下游任务。其临床应用面临两大挑战:(i) 协变量偏移——因训练数据分散于不同医院、设备和时间,导致潜在动态预测器所见特征分布与部署时存在差异;(ii) 置信度错配——多步预测在临床风险最高处往往过度自信。本文提出统一解决方案CalTwin,结合源自先前工作的费雪信息偏移惩罚(用于处理碎片化协变量偏移)与置信度错配惩罚(用于校准视觉-语言分类),应用于基于GRU的医疗世界模型的潜在状态转移预测器。我们推导了联合目标函数,明确哪些证明步骤可直接迁移,哪些需调整,并在PhysioNet 2019脓毒症挑战赛上进行评估,将两个医院系统视为顺序训练片段,未见系统作为分布外测试。相比无惩罚基线,CalTwin使外部测试集下一步潜态均方误差降低9.1%(仅费雪信息惩罚贡献7.0%);置信度错配惩罚带来的期望校准误差(ECE)下降为0.7%(仅该惩罚项达1.3%),效果真实但较小。
原文摘要 · Abstract (English)
Medical world models aim to learn a latent state of patient or organ physiology and a transition function that forecasts how that state evolves under interventions, supporting downstream tasks from imaging-based diagnosis to digital-twin treatment planning. Two failure modes threaten the reliability of such models in clinical deployment: (i)~\emph{covariate shift}, because training data are fragmented across hospitals, scanners, and time, so the feature distribution seen by the latent-dynamics predictor differs across fragments and from the distribution at deployment; and (ii)~\emph{confidence misalignment}, because multi-step forecasts are often overconfident exactly where clinical risk is highest. We argue that both problems admit a unified treatment via a single lightweight regularisation objective, \textbf{CalTwin}, which combines a Fisher-Information-based shift penalty adapted from our prior work on fragmented covariate-shift remediation~\cite{khan2025mitigating,khan2025causal} with a Confidence Misalignment Penalty adapted from our prior work on calibrated vision-language classification~\cite{khan2025confidence}, applied here to a GRU-based medical world model's latent transition predictor. We derive the combined objective, establish which proof steps transfer from the classification setting without modification and which require adaptation, and evaluate it on the PhysioNet 2019 Sepsis Challenge, treating the two hospital systems as sequential training fragments and the unseen system as an out-of-distribution test. CalTwin reduces OOD next-step latent-state MSE by 9.1\% relative to the no-penalty baseline (FIM penalty alone accounts for 7.0\%); the ECE reduction from the Confidence Misalignment Penalty is real but small (0.7\% for CalTwin, 1.3\% for CMP alone).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。