用多模态模型从CT和病历中推断患者隐状态,实现诊断、重建与模拟。
HounsWorld: A Multimodal World Model for Hidden Patient-State Readout, Reconstruction, and Simulation

- 构建共享患者状态的多模态世界模型,融合CT图像与临床语言
- 在3个任务上均表现优异,零样本下仍能完成低剂量降噪与虚拟增强
- 适合医学影像智能、临床决策支持方向的研究者使用
临床智能需从不完整观测中推断患者的潜在状态,而非仅学习影像到答案的孤立映射。体积分层医学图像提供解剖结构、衰减和病灶的密集观测,而临床文本则提供稀疏但互补的语义信息。我们提出以CT为中心的智能建模,将读取、重建与仿真统一为对共享隐式患者状态的条件预测。为此,我们构建了HounsBench——一个以CT为中心的患者状态基准,涵盖三类任务,采用患者互斥划分和每类独立评估指标;并提出HounsWorld,一个30亿参数的多模态世界模型,通过联合理解-生成学习将体积分层扫描与语言视为共享状态的观测。共享Transformer隐式估计患者状态,支持三种输出:基于查询的答案读出、语言形式的报告与描述重建,以及条件化生成的低剂量去噪CT、虚拟对比增强图像、及文本/掩码驱动的解剖约束体积生成。零初始化的CT适配器保留预训练的多模态映射,条件显式的亨氏单位窗口采样则揭示具有临床意义的密度信息。HounsWorld在三类任务中均表现强劲,并通过结构化补全持续提升对CT的理解。项目开源地址:https://github.com/byhwhite/HounsWorld.git
原文摘要 · Abstract (English)
Clinical intelligence requires estimating a patient's underlying condition from incomplete observations rather than learning isolated mappings from scans to answers. Volumetric medical images provide dense observations of anatomy, attenuation, and lesions, whereas clinical language provides sparse but complementary semantic observations. We formulate CT-centered intelligence as inference over a shared latent patient state, under which readout, reconstruction, and simulation all become state-dependent prediction problems. To operationalize this view, we introduce HounsBench, a computed tomography (CT) centric patient-state benchmark that unifies these three task families with patient-disjoint splits and per-family metrics, and HounsWorld, a 3B multimodal world model that treats volumetric scans and language as observations of the shared state through Joint Understanding-Generation Learning. A shared transformer forms an implicit patient-state estimate and supports three outputs: query-conditioned answers that read out the state, reports and captions that reconstruct it in language, and condition-specific CT volumes for low-dose denoising, virtual contrast enhancement, and anatomy-constrained text-and-mask-to-volume generation. Zero-initialized CT adapters preserve pretrained multimodal mappings, while condition-explicit Hounsfield-unit window sampling exposes clinically meaningful density observations. HounsWorld shows strong performance across all three task families while consistently improving CT understanding through clinically structured completion. Our project is available at https://github.com/byhwhite/HounsWorld.git
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。