用视觉语言模型自动识别胎儿脑区细胞结构,提升神经解剖分析效率。
CytoCLIP: Learning Cytoarchitectural Characteristics in Developing Human Brain Using Contrastive Language Image Pre-Training
- 基于CLIP框架构建双尺度模型,分别学习整体区域与细胞级结构特征。
- 在86个大区域和379个细粒度区域上实现0.87和0.91的分类准确率。
- 适用于发育中人脑研究,助力自动化脑区标注与跨模态检索。
人类大脑不同区域的功能与其独特的细胞架构密切相关,细胞的空间分布与形态定义了其细胞架构。通过细胞架构识别脑区可支持多种脑科学研究。然而,手动标注脑组织切片中的脑区耗时且需专业知识。为此,我们提出CytoCLIP,一套基于对比语言-图像预训练(CLIP)框架的视觉语言模型,用于学习脑细胞架构的联合视觉-文本表征。该模型包含两种变体:一种使用低分辨率全区域图像,捕捉区域整体细胞架构模式;另一种使用高分辨率图像块,实现细胞级别表征。训练数据来自不同孕周发育中胎儿脑的尼氏染色组织切片,包含86个低分辨率区域和379个高分辨率图像块。通过区域分类与跨模态检索任务评估模型对细胞架构的理解能力与泛化性能。在不同样本年龄与切片平面的数据设置下进行多组实验,结果表明,CytoCLIP优于现有方法,在全区域分类上达到0.87的加权F1分数,在高分辨率图像块分类上达0.91。
原文摘要 · Abstract (English)
The functions of different regions of the human brain are closely linked to their distinct cytoarchitecture, which is defined by the spatial arrangement and morphology of the cells. Identifying brain regions by their cytoarchitecture enables various scientific analyses of the brain. However, delineating these areas manually in brain histological sections is time-consuming and requires specialized knowledge. An automated approach is necessary to minimize the effort needed from human experts. To address this, we propose CytoCLIP, a suite of vision-language models derived from pre-trained Contrastive Language-Image Pre-Training (CLIP) frameworks to learn joint visual-text representations of brain cytoarchitecture. CytoCLIP comprises two model variants: one is trained using low-resolution whole-region images to understand the overall cytoarchitectural pattern of an area, and the other is trained on high-resolution image tiles for detailed cellular-level representation. The training dataset is created from NISSL-stained histological sections of developing fetal brains of different gestational weeks. It includes 86 distinct regions for low-resolution images and 379 brain regions for high-resolution tiles. We evaluate the model's understanding of the cytoarchitecture and generalization ability using region classification and cross-modal retrieval tasks. Multiple experiments are performed under various data setups, including data from samples of different ages and sectioning planes. Experimental results demonstrate that CytoCLIP outperforms existing methods. It achieves a weighted F1 score of 0.87 for whole-region classification and 0.91 for high-resolution image tile classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。