无需修改模型,通过局部提示揭示视觉大模型的隐藏空间结构。
P3CA: Encoder-Agnostic Interpretation of Vision Foundation Model Embeddings via Spatial Probing

- 用空间提示引导主成分分析,定位关键特征方向。
- 在病理图像中提升提示匹配的分类准确率,优于全局PCA。
- 适用于医学影像、基因表达等多模态数据的可解释性分析。
视觉基础模型在医学图像计算中被广泛用作可复用编码器,但其高维空间嵌入难以在下游任务性能之外进行直观解析。我们提出位置提示主成分分析(P3CA),一种无需依赖编码器的方法,用于对通道丰富的空间张量进行局部探查。给定用户选定的空间提示,P3CA估计该区域内特征的归一化和主导协方差方向,并将所得投影应用于整个张量,以可视化局部信息丰富方向的分布。该方法不需修改编码器、重新训练或任务标签,即可生成区域条件化的表示视角。我们在EmbedVision中实现P3CA,这是一个基于3D Slicer的交互式工作流,评估了其在自然图像、结直肠病理基础模型嵌入及空间转录组张量中的表现。结果表明,提示驱动的投影能揭示全局PCA所掩盖的局部结构,在冻结的三维投影中提升病理判别能力,并支持学习到的与实际测量的空间表示之间的对比。
原文摘要 · Abstract (English)
Vision foundation models are increasingly used as reusable encoders in medical image computing, yet their high-dimensional spatial embeddings are difficult to inspect beyond downstream task performance or global dimensionality reduction. We propose position-prompted PCA (P3CA), an encoder-agnostic method for local probing of channel-rich spatial tensors. Given a user-selected spatial prompt, P3CA estimates the feature normalization and dominant covariance directions within that region, then applies the resulting projection to the full tensor to visualize where locally informative directions are expressed. This produces a region-conditioned representation lens without modifying the encoder, retraining, or requiring task-specific labels. We implement P3CA in EmbedVision, an interactive 3D Slicer-based workflow, and evaluate it across natural images, colorectal pathology foundation-model embeddings, and spatial transcriptomic tensors. Across these settings, prompted projections reveal local structure suppressed by global PCA, improve prompt-matched pathology discrimination from frozen three-dimensional projections, and support comparison between learned and measured spatial representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。