用文献文本弱监督训练脑组织图像语言模型,实现精准描述与信息读取。
Cytoarchitecture in Words: Weakly Supervised Vision-Language Modeling for Human Brain Microscopy
- 从文献中检索标签文本并结合图像特征增强,构建弱监督训练数据
- 在57个脑区上实现90.6%的区域识别准确率,描述仍可支持8分类识别
- 无需显式区域名也能还原关键结构信息,适合神经解剖学研究者使用
视觉基础模型正推动科学工作流的交互化,但自然语言交互需视觉与语言表征耦合。在人类大脑细胞体染色组织切片中,显微图像块编码了细胞密度、形态、分层结构和区域组织等皮质架构信息。我们提出“检索-增强”监督机制,无需人工标注图文对即可训练图像条件语言模型:通过共享解剖标签从文献中检索文本,再结合皮质层厚度、细胞密度等图像特异性属性进行增强。标签仅作训练目标,不作为模型输入。利用该方法,我们将CytoNet(皮质架构视觉基础模型)与开源大语言模型通过轻量级Flamingo风格适配器耦合。在57个脑区上,模型生成合理皮质架构描述,支持开放集使用(拒绝非目标区域),对目标区域图像的预测准确率达90.6%。即使移除生成文本中的区域名称,语言模型仍能在8分类任务中以68.6%准确率恢复区域信息。另一经指令微调的模型能从单个图像块中读出皮质层厚度与细胞密度,实现超越标准区域描述的信息提取。结果表明,该方法为缺乏图像级标注但存在标签与专家文本的专业成像领域提供了实用的视觉-语言训练路径。
原文摘要 · Abstract (English)
Vision foundation models increasingly support interactive scientific workflows, but natural-language interaction requires coupling visual representations to language. Curated image-text pairs for this coupling are scarce in many biomedical domains. This is the case for cell-body-stained histological sections of the human brain, where microscopic image patches encode cytoarchitecture: cellular density, morphology, laminar structure, and areal organization. We propose retrieve-and-enrich supervision, a weakly supervised scheme for training image-conditioned language models without curated image-text pairs. The method retrieves label-level text from the literature via a shared anatomical label, then enriches it with image-specific properties such as cortical layer thickness and cell density. Labels provide training targets only, not model inputs. We use this scheme to couple CytoNet, a cytoarchitectonic vision foundation model, to an open-weight large language model via a lightweight Flamingo-style adapter. Across 57 brain areas, the resulting model produces plausible cytoarchitectonic descriptions, supports open-set use by rejecting out-of-scope areas, and predicts the correct area for in-scope patches with 90.6% accuracy. Removing explicit area names from generated text still leaves descriptions sufficient for a language model to recover the area in an 8-way test with 68.6% accuracy. A second instruction-tuned model trained with image-specific targets recovers cortical layer thickness and cell-densities from individual patches, allowing the model to read out information beyond canonical area descriptions. Our results show that retrieve-and-enrich supervision offers a practical route to vision-language training in specialized imaging domains where labels and expert text exist but image-level captions do not.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。