arXiv:2412.10659cs.CVcs.LG2024-12AAAI被引 7

融合组织图像与空间转录组数据,提升细微异常组织区域检测精度

MEATRD: Multimodal Anomalous Tissue Region Detection Enhanced with Spatial Transcriptomics

  • 通过多模态重构与单类分类结合,利用图像和基因表达数据联合识别异常区域
  • 在8个真实空间转录组数据集上超越现有先进方法,对微弱视觉差异的异常区域仍有效
  • 提出新型掩码图双注意力网络,解决重建类方法过拟合问题,首次分析多模态瓶颈编码信息特性

异常组织区域(ATRs)的检测在临床诊断和病理研究中至关重要。传统基于组织学图像的方法在异常区域与正常组织视觉差异微小时表现不佳。空间转录组(ST)技术可表征组织区域的基因表达,为检测ATRs提供分子视角。然而,尚缺乏有效整合图像与ST数据的ATR检测方法。为此,我们提出MEATRD,一种融合组织学图像与ST数据的新型检测方法。MEATRD通过重建正常组织点(内点)的图像块和基因表达谱,学习基于多模态嵌入的单类异常检测模型,结合了重建与单类分类的优势。核心是创新的掩码图双注意力变换器(MGDAT)网络,实现跨模态、跨节点信息共享,并缓解重建类方法的过泛化问题。此外,我们首次揭示多模态瓶颈编码能凝聚模态特异且任务相关的有效信息。在八个真实ST数据集上的大量评估显示,MEATRD在ATR检测上优于多种先进方法,尤其擅长识别仅具轻微视觉偏差的异常区域。

原文摘要 · Abstract (English)

The detection of anomalous tissue regions (ATRs) within affected tissues is crucial in clinical diagnosis and pathological studies. Conventional automated ATR detection methods, primarily based on histology images alone, falter in cases where ATRs and normal tissues have subtle visual differences. The recent spatial transcriptomics (ST) technology profiles gene expressions across tissue regions, offering a molecular perspective for detecting ATRs. However, there is a dearth of ATR detection methods that effectively harness complementary information from both histology images and ST. To address this gap, we propose MEATRD, a novel ATR detection method that integrates histology image and ST data. MEATRD is trained to reconstruct image patches and gene expression profiles of normal tissue spots (inliers) from their multimodal embeddings, followed by learning a one-class classification AD model based on latent multimodal reconstruction errors. This strategy harmonizes the strengths of reconstruction-based and one-class classification approaches. At the heart of MEATRD is an innovative masked graph dual-attention transformer (MGDAT) network, which not only facilitates cross-modality and cross-node information sharing but also addresses the model over-generalization issue commonly seen in reconstruction-based AD methods. Additionally, we demonstrate that modality-specific, task-relevant information is collated and condensed in multimodal bottleneck encoding generated in MGDAT, marking the first theoretical analysis of the informational properties of multimodal bottleneck encoding. Extensive evaluations across eight real ST datasets reveal MEATRD's superior performance in ATR detection, surpassing various state-of-the-art AD methods. Remarkably, MEATRD also proves adept at discerning ATRs that only show slight visual deviations from normal tissues.

异常检测空间转录组多模态融合医学图像分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。