冷冻脑影像模型的嵌入其实记录了扫描设备信息,可能影响研究结果。
Frozen Brain-MRI Foundation Models Are Site Fingerprints

- 用线性分类器可高精度识别扫描站点,即使随机初始化模型也能做到
- 站点信息在深层特征中占比达0.9平衡准确率,远超性别、年龄等临床变量
- 该指纹源于图像低层统计特征,建议使用前进行站点审计
冻结的基础模型(FM)嵌入被广泛用作现成的脑影像表示,假设其捕捉的是解剖结构。我们审计了这些嵌入实际编码的内容,发现采集站点是表征中的主要内在成分。在两个独立队列(ABIDE-I, ABIDE-II)中,三种冻结的3D编码器(脑预训练、CT预训练、随机初始化)及所有网络深度下,站点在线性可解度上均达到约0.9的平衡准确率,超过所有临床或人口学变量(性别、年龄、自闭症诊断)在各层的表现。该效应是内在而非学习所得:随机初始化编码器在两个队列和三种架构(Swin、ViT、ResNet)下已具备约0.9的站点分类能力;甚至直接从降采样原始图像中即可实现约0.95的站点解码,说明指纹反映的是任何编码器都会保留的低层图像统计特征,而非预训练产物。残差化人群协变量后,站点可解度基本不变,表明其为采集驱动而非人群驱动。非线性探测器与线性探测器性能相当,说明该指纹完全线性可访问。通过迭代零空间投影或ComBat可移除站点子空间(解码率从0.94降至0.07/0.00),但用于密集分割时代价高昂,因站点与解剖结构占据纠缠的线性子空间(匹配秩随机方向投影保持Dice值不变,而移除站点子空间则破坏解剖信息)。建议对冻结脑影像基础模型进行站点审计,并开源审计工具包。
原文摘要 · Abstract (English)
Frozen foundation-model (FM) embeddings are increasingly used as off-the-shelf brain-MRI representations, on the assumption that they capture anatomy. We audit what they actually encode and find that acquisition site is a large, intrinsic component of the representation. Across two independent cohorts (ABIDE-I, ABIDE-II), three frozen 3-D encoders (brain-pretrained, CT-pretrained, and randomly initialized), and every network depth, site is linearly decodable at roughly 0.9 balanced accuracy at deep layers, exceeding the decodability of every clinical or demographic variable (sex, age, autism diagnosis) at every layer. The effect is intrinsic rather than learned: a randomly initialized encoder is already a ~0.9 site classifier on both cohorts and across three architecture families (Swin, ViT, ResNet), and site is decodable at ~0.95 directly from the raw downsampled image with no encoder, so the fingerprint reflects low-level image statistics that any encoder preserves rather than a product of pretraining. Residualizing measured population covariates leaves site decodability essentially unchanged, indicating an acquisition- rather than population-driven effect. A nonlinear probe matches the linear one, so the fingerprint is fully linearly accessible. The site subspace is removable post hoc by iterative null-space projection or ComBat (site decodability 0.94 -> 0.07/0.00), and is a site-attribution concern for shared or federated embeddings; but for dense segmentation this removal is not free, because site and anatomy occupy an entangled linear subspace (a matched-rank random-direction projection is Dice-neutral, whereas removing the site subspace is destructive). We recommend site-audited use of frozen brain-MRI FMs and release an open audit toolkit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。