提出结构优先的病理图像生成框架,提升跨模态一致性与可控性。
PathAR: Structure-First Autoregressive Synthesis of Multimodal Pathology Images

- 分离结构与外观特征,通过双向量化编码实现解耦建模。
- 在多模态病理数据上显著提升结构一致性和模态保真度。
- 适合医学图像生成、数据稀缺场景下的下游分割任务。
多模态病理数据稀缺推动了统一生成模型的发展,要求在保留解剖结构一致性的同时合成特定模态的外观。尽管不同模态在外观统计上存在差异,但细胞拓扑和组织边界等形态结构在多种采集协议下基本保持不变。然而,现有方法常将这些因素混合于同质令牌流中,隐式耦合结构与外观,导致在模态变换下结构控制能力减弱。为此,我们提出病理自回归建模(PathAR),一种结构优先的自回归合成框架,显式分解结构与外观,实现模态标签条件下的病理图像生成。PathAR采用基于掩码的双向量量化(Dual-VQ)分词器,将样本分解为掩码引导的结构与外观令牌,并使用具有非对称注意力可见性的交错自回归(IAR)Transformer,强制结构到外观的依赖关系。实验表明,PathAR在异构模态外观下稳定保持形态一致性,支持空间对齐的图像-掩码对生成,在基准测试中优于现有方法,保持样本多样性,并在数据稀缺环境下支持下游分割任务,且可扩展至更细粒度的模态内器官标签变化。
原文摘要 · Abstract (English)
Data scarcity in multimodal pathology motivates unified generative models that synthesize modality-specific appearance while preserving anatomically coherent structure. Although modalities differ in appearance statistics, morphological structures such as cellular topology and tissue boundaries are largely preserved across acquisition protocols. However, existing methods often model these factors within a homogeneous token stream, implicitly coupling structure with appearance and weakening structural controllability under modality shifts. To address this, we propose pathology Autorgressive modeling (PathAR), a structure-first autoregressive synthesis framework that explicitly factorizes structure and appearance for modality-label-conditioned pathology generation.PathAR employs a dual vector quantization (Dual-VQ) tokenizer to decompose samples into mask-grounded structure and appearance tokens, and an interleaved autoregressive (IAR) transformer with asymmetric attention visibility to enforce structure-to-appearance dependence. PathAR stabilizes morphology under heterogeneous modality-specific appearances and enables spatially aligned image--mask pair generation. Extensive experiments show that PathAR improves structural consistency and modality fidelity over baselines, maintains sample diversity, supports downstream segmentation in data-scarce regimes, and demonstrates extensibility to finer-grained intra-modality organ-label variation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。