arXiv:2606.20390cs.CV2026-06中稿 · MICCAI 2026

用图结构建模皮肤病变区域,融合元数据提升分类准确率

Geometry-Aware Superpixel Graph Transformer with Metadata for Skin Lesion Classification

论文配图:Geometry-Aware Superpixel Graph Transformer with Metadata for Skin Lesion Classification
图 1 · 摘自论文原文
  • 将病变区域抽象为超像素图节点,通过几何关系编码区域间空间结构
  • 引入元数据节点与所有区域连接,在同一图空间内融合临床信息,准确率提升显著
  • 适合需要多模态融合的医学图像分类任务,尤其关注细粒度结构分析

基于皮肤镜图像的皮肤癌自动分类仍面临病灶结构异质、类内差异大以及良恶性之间视觉差异微弱等挑战。现有CNN/ViT方法通常依赖全局或块级特征,并通过后期融合方式结合患者元数据,限制了空间感知的多模态推理。本文提出一种基于区域的图学习框架,将病变显式建模为由空间连贯的超像素区域构成的图,每个区域以冻结的CNN特征表示。为捕捉精细的病变排列模式,通过边属性编码区域间的几何关系,并引入一个专属的元数据上下文节点,与所有区域相连,实现人口统计学/临床变量在统一关系空间中的结构化融合。使用边感知图变压器更新节点表示,并通过注意力驱动传播,最终生成图级嵌入用于良恶性分类。在四个公开基准上的实验表明,显式的区域级关系建模与原生多模态融合带来持续性能提升,超越当前最优方法。由此确立了一种新的图中心视角:将CNN特征视为关系节点,通过上下文整合获得更丰富、更鲁棒的表征。

原文摘要 · Abstract (English)

Automated skin cancer classification from dermoscopic images remains challenging due to heterogeneous lesion structure, strong intra-class variability, and subtle visual differences between benign and malignant cases. Existing CNN/ViT pipelines typically rely on global or patch-level features and often combine patient metadata via late fusion, which limits spatially grounded multimodal reasoning. We present a novel region-based graph learning framework that explicitly models lesions as graphs of spatially coherent superpixel regions represented as frozen CNN features. To capture fine-grained lesion arrangements, we encode inter-regional geometry as edge attributes and introduce a dedicated metadata context node connected to all regions, providing structured integration of demographic/clinical variables within the same relational space. Node representations are updated using our edge-aware graph transformer followed by attention-driven propagation, and a final graph-level embedding for benign-malignant classification. Experiments on four public benchmarks demonstrate that explicit region-level relational modeling and graph-native multimodal fusion yield consistent gains over the state-of-the-art. Consequently, we establish a new graph-centric perspective in which CNN features are modeled as relational nodes and improved through contextual integration, yielding more expressive and robust classifications.

皮肤病变图神经网络多模态融合医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。