3D自编码器更擅长学习细胞显微图像的立体特征,提升定位与互作预测性能。
3D Masked Autoencoders are Robust Learners of Volumetric and Multimodal Cellular Representations for Microscopy

- 用3D掩码自编码器直接处理体数据,比2D投影更充分挖掘空间信息
- 在蛋白定位任务上达到0.95的AUC,互作预测达0.86的ROC-AUC
- 结合蛋白质语言模型可进一步提升性能,适合单细胞多模态研究者
荧光显微镜中的自监督学习常依赖二维投影,但细胞本身具有三维结构。本文系统比较了2D与3D掩码自编码器(MAE-2D vs. MAE-3D)在体数据上的表现。在相同架构与训练条件下,MAE-3D在下游单细胞任务中持续优于2D最大投影和切片变体。通过将视觉表示与预训练蛋白语言模型ESM2对齐,发现跨模态监督对体数据模型收益更大。通道交叉注意力和频域正则化对利用3D空间上下文至关重要。在蛋白-蛋白互作预测任务中,最佳模型达到ROC-AUC 0.86;在蛋白定位任务中,AUC_micro为0.95,F1_micro为0.74,表现优异。结果表明,体建模与多模态对齐在单细胞显微表征学习中潜力巨大。
原文摘要 · Abstract (English)
Self-supervised learning in fluorescence microscopy often relies on 2D projections, despite the inherently three-dimensional nature of cells. We present a systematic comparison of 2D and 3D masked autoencoders (MAE-2D vs. MAE-3D) on volumetric microscopy data. Under matched architectures and training protocols, MAE-3D consistently outperforms 2D max-projection and slice-based variants on downstream single-cell tasks. We further align visual representations with a pretrained protein language model (ESM2) and show that cross-modal supervision yields larger gains for volumetric models. Channel cross-attention and frequency-domain regularization are critical for leveraging 3D spatial context. On protein--protein interaction prediction, our best model achieves a ROC--AUC of 0.86, while on protein localization it reaches an AUC$_{\text{micro}}$ of 0.95 and an F1$_{\text{micro}}$ of 0.74, demonstrating competitive performance on both tasks. Overall, our findings highlight the potential of volumetric modeling and multimodal alignment for representation learning in single-cell microscopy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。