arXiv:2609.07719cs.AI2026-09

构建可同时诊断与生成影像的医学影像世界模型。

A radiographic world model for clinical reasoning and evidence generation

论文配图:A radiographic world model for clinical reasoning and evidence generation
图 1 · 摘自论文原文
  • 基于265万对胸片-文本数据,学习共享连续潜在状态。
  • 诊断准确率超现有模型,合成数据提升真实数据表现。
  • 可针对性生成特定人群证据,改善公平性问题。

医学影像人工智能通常独立训练从影像到诊断或从描述到图像的映射,但二者均源自同一影像状态。本文提出MedDream,一种放射科世界模型,通过265万对去泄漏控制的胸片-文本配对数据(来自440万候选)预训练,学习共享的连续潜在状态,支持诊断推理与报告条件下的影像生成。在八个临床数据集和两组独立阅片者中,MedDream优于主流诊断与生成模型。诊断方面,其在疾病识别、少标签适应、严重程度评估和定位上均表现优异,且辅助阅片使住院医师与独立放射科医生共识一致性从56.3%提升至63.0%。生成方面,合成影像保留关键病理特征,提升下游任务性能:在外部VinDr-CXR数据集上,合成增强使宏平均AUROC从76.4%提升至81.4%。更重要的是,基于预设亚组性能差距进行条件生成,使亚洲患者加权F1提升3.1个百分点,而等量无指导增强反而下降2.3个百分点。结果表明,放射科世界模型是实现可解释、可模拟、可构造证据的临床可用医学AI的关键路径。

原文摘要 · Abstract (English)

Medical imaging artificial intelligence (AI) is commonly developed as separate mappings from radiographs to diagnostic outputs or from clinical descriptions to generated images, although both arise from the same underlying radiographic state. A world-model formulation instead seeks to learn an internal representation of this state that can support both clinical readout and conditional simulation of radiographic observations. Here we introduce MedDream, a radiographic world model that learns a shared continuous latent state from paired chest radiograph-text observations for diagnostic reasoning and report-conditioned evidence generation. MedDream was pretrained on 2.65 million leakage-controlled chest radiograph-text pairs curated from 4.40 million candidates. Across eight clinical datasets and two independent reader cohorts, MedDream outperformed leading diagnostic and generative comparators. For diagnostic reasoning, MedDream showed strong generalization across disease recognition, label-scarce adaptation, severity assessment, and localization, while MedDream-supported review increased mean resident concordance with independent radiologist consensus from 56.3% to 63.0%. For evidence generation, MedDream produced radiographs that preserved clinically relevant pathology and improved downstream performance on held-out real data, with synthetic augmentation increasing external VinDr-CXR macro-AUROC from 76.4% to 81.4%. More importantly, conditioning generation on prespecified subgroup performance gaps enabled targeted evidence construction, increasing weighted F1 by 3.1 percentage points in Asian patients, whereas matched-volume unguided augmentation decreased it by 2.3 points. These findings establish radiographic world models as a path toward medical AI that learns clinically meaningful internal states for interpreting, simulating, and constructing evidence for clinical use.

医学影像世界模型生成诊断公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。