arXiv:2607.20274cs.CVcs.AI2026-07

自监督训练比临床标注更能推动医学模型表示收敛。

Self-supervision drives representational convergence in medical foundation models more than clinical supervision

论文配图:Self-supervision drives representational convergence in medical foundation models more than clinical supervision
图 1 · 摘自论文原文
  • 控制变量实验发现,自监督目标主导表示收敛,而非临床标签或模型规模。
  • 自监督模型在胸片任务上对齐率达40.4%,远超临床监督的21.1%。
  • 模型间可迁移性良好,适合跨机构医疗系统部署与验证。

不同研究团队的医学图像编码器日益被视为可互换,假设其表征因规模和临床监督趋于一致。但这种收敛是否真实、由何驱动、是否具备临床价值尚未验证,现有相似性度量也脆弱。本文通过18个图像与7个文本编码器的受控剖析,涵盖700万至270亿参数、五种成像模态及65万张胸片数据(来自六个数据集),在固定数据、架构和规模下,仅改变训练目标。结果表明:表示收敛程度有限但高于随机基线,主要由自监督目标驱动,非临床监督;相同自监督编码器在胸片任务上对齐率高达40.4%,而标签监督仅为21.1%,图文对比仅3.3%;收敛不随模型大小或能力增长(斯皮尔曼相关系数0.302,p=0.223)。该收敛为同模态内现象,无法达到临床语言水平,且无法复现放射科医生判断病例相似性的模式。然而,线性分类器可在编码器间迁移,并在五个独立医院中保持约85%的原模型性能。因此,医学影像中的表示收敛取决于预训练目标,而非规模或临床监督。互操作性应通过目标设计实现,并在患者子群体和临床判断下重点验证共享几何结构的薄弱环节。

原文摘要 · Abstract (English)

Medical image encoders from different groups are increasingly treated as interchangeable, on the assumption that scale and clinical supervision concentrate their representations onto a shared structure. Whether this convergence is real, what produces it, and whether it is clinically usable are untested, and the similarity measures behind such claims are fragile. We present a controlled dissection across 18 image and 7 text encoders, all open-weight and run locally, spanning 7M to 27B parameters and five imaging modalities, including 650,982 chest radiographs from six datasets. To isolate cause, we train encoders that vary only the objective under fixed data, architecture, and scale, and reproduce the effect in a synthetic model. Convergence is modest but above a random floor, driven by the self-supervised objective, not clinical supervision: matched self-supervised encoders aligned most (40.4% on chest radiography), with label-supervised (21.1%) and image-text (3.3%) far lower, and did not grow with size (Spearman 0.302, p=0.223) or capability. It is within-modality, does not reach clinical language, and does not reproduce how radiologists judge case similarity. Yet a linear classifier transfers across encoders and to five held-out hospitals, retaining about 85% of within-encoder performance. Convergence in medical imaging is therefore set by the pretraining objective, not inherited from scale or clinical supervision. Interoperability is accordingly something to design for through that objective, and to validate where the shared geometry is weakest, across patient subgroups and against clinical judgment.

医学模型自监督表示学习可迁移性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。