用轻量级Transformer解析多重组织图像中的细胞空间结构
MUL-T: Decoding Spatial Cellular Architecture in Multiplexed Tissue Images

- 将组织架构建模为细胞标记的上下文预测任务
- 在多个临床任务中表现优于传统方法,接近大模型性能
- 适合需要高效分析多标记组织图像的研究者
解析多重组织成像中的组织结构需同时建模细胞表型及其空间关系。现有方法通常依赖手工特征(如标记强度统计或细胞类型比例),难以跨不同标记组合的队列推广。我们提出MUL-T,一种轻量级Transformer框架,将组织架构重构为离散细胞标记上的掩码上下文预测任务。通过无任务特定监督学习上下文嵌入[CLS],模型捕捉高阶细胞交互,同时保持计算高效。我们在多个临床相关下游任务中评估MUL-T,包括核心级别肿瘤模式分类、患者级别分级、PD-L1阳性预测及跨数据集治疗反应预测。在各项任务中,MUL-T均持续优于经典特征基线,并达到与基础ViT模型相当的性能,参数量更少,训练成本更低。
原文摘要 · Abstract (English)
Understanding tissue organisation in multiplexed imaging requires modelling both cellular phenotypes and their spatial context. Existing approaches typically rely on handcrafted features, such as marker intensity statistics or cell-type proportions, which often fail to scale or generalise across cohorts with heterogeneous marker panels. We introduce MUL-T, a lightweight transformer framework that reframes tissue architecture as a masked contextual prediction task over discrete cell tokens. By learning contextualised [CLS] embeddings without task-specific supervision, the model captures higher-order cellular interactions while remaining computationally efficient. We evaluate MUL-T on several clinically relevant downstream tasks, including core-level tumour pattern classification, patient-level grading, PD-L1 positivity prediction, and cross-dataset treatment response prediction. Across tasks, MUL-T consistently outperforms classical feature-based baselines and achieves performance comparable to a foundation ViT model, despite substantially fewer parameters and lower training cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。