arXiv:2605.26026cs.CVcs.AI2026-05被引 1

基于多模态3D基础模型,实现光片荧光显微镜数据的少样本分割与去模糊。

A Multimodal 3D Foundation Model for Light Sheet Fluorescence Microscopy Enables Few-Shot Segmentation, Classification, and Deblurring

论文配图:A Multimodal 3D Foundation Model for Light Sheet Fluorescence Microscopy Enables Few-Shot Segmentation, Classification, and Deblurring
图 1 · 摘自论文原文
  • 联合优化掩码重建与图文对齐,学习可迁移的3D体积表征。
  • 少样本下在分割、分类和去模糊任务上均超越基线模型。
  • 适合生物图像分析、需减少人工标注的研究者使用。

光片荧光显微镜(LSM)可实现高分辨率三维生物样品成像,提供丰富的体积数据用于研究细胞结构、病理变化及血管网络。然而,LSM数据规模大、维度高且标注成本高,使监督深度学习方法难以扩展。尽管存在大量未标注的LSM体数据,但该模态的基础模型仍因计算挑战和体积表示学习复杂性而研究不足。本文提出一种针对LSM数据的3D基础模型,基于涵盖多种生物体、染色方式和成像协议的大规模预训练数据集进行训练。通过联合优化掩码重建与图像-文本对齐,学习可迁移的体积表示。预训练主干网络显著降低标注需求,支持高效少样本适配多种下游任务。我们在分割、分类和去模糊任务上评估该方法,结果表明其在标准指标和领域专家严格评估中均持续优于基线。这证明了基础模型预训练能有效降低标注要求并提升多样化的LSM分析性能。预训练权重及代码已公开:https://github.com/AdinaScheinfeld/lsm_fm_public_repo.git。

原文摘要 · Abstract (English)

Light sheet fluorescence microscopy (LSM) enables high-resolution, three-dimensional (3D) imaging of biological specimens, providing rich volumetric data for studying cellular organization, pathology, and vascular networks. However, the size, dimensionality, and annotation burden of LSM data make supervised deep learning approaches costly and difficult to scale. Additionally, despite the abundance of unannotated LSM volumes, foundation models for this modality remain underexplored due to computational challenges and the complexity of volumetric representation learning. In this work, we introduce a 3D foundation model for LSM data, pretrained on a large curated collection of 3D images spanning multiple organisms, stains, and imaging protocols. We learn transferable volumetric representations by jointly optimizing for masked reconstruction and image-text alignment. The pretrained backbone drastically reduces the annotation burden, enabling efficient, few-shot adaptation for varied downstream tasks. We evaluate this approach on downstream segmentation, classification, and deblurring. Our results demonstrate consistent improvements over baselines, (1) when measured using standard evaluation metrics and (2) when rigorously assessed by domain experts. This highlights the potential of foundation model pretraining to reduce annotation requirements while improving performance across diverse LSM analysis tasks. Pretrained model weights and code for pretraining and finetuning are publicly available: https://github.com/AdinaScheinfeld/lsm_fm_public_repo.git.

3D建模生物图像基础模型少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。