用布局引导生成病理图像,精准控制组织结构和细节。
Layout-Guided Controllable Pathology Image Generation with In-Context Diffusion Transformers
- 引入上下文扩散变压器,融合布局、文本与视觉特征
- 在5个数据集上实现更高保真度与空间可控性
- 适合医学影像生成与数据增强研究者使用
可控病理图像合成需精确调控空间布局、组织形态与语义细节。现有文本引导的扩散模型仅提供粗粒度全局控制,缺乏细粒度结构约束能力。进展受限于缺乏大规模配对数据集——将切片级空间布局与详细诊断描述配对,对人类专家而言在千兆像素全切片图像上标注耗时极长。为此,我们构建了一套可扩展的多智能体LVLM标注框架,整合图像描述生成、诊断步骤提取与自动质量判断,通过人工验证评估系统可靠性,实现大规模细粒度、临床对齐的监督数据构建。基于此数据,提出上下文扩散变压器(IC-DiT),一种布局感知生成模型,将空间布局、文本描述与视觉嵌入统一输入扩散变压器。通过分层多模态注意力机制,IC-DiT在保持全局语义一致性的同时,精准保留结构与形态细节。在五个组织病理学数据集上的实验表明,IC-DiT在保真度、空间可控性与诊断一致性方面优于现有方法。生成图像还可作为癌症分类与生存分析等下游任务的有效数据增强资源。
原文摘要 · Abstract (English)
Controllable pathology image synthesis requires reliable regulation of spatial layout, tissue morphology, and semantic detail. However, existing text-guided diffusion models offer only coarse global control and lack the ability to enforce fine-grained structural constraints. Progress is further limited by the absence of large datasets that pair patch-level spatial layouts with detailed diagnostic descriptions, since generating such annotations for gigapixel whole-slide images is prohibitively time-consuming for human experts. To overcome these challenges, we first develop a scalable multi-agent LVLM annotation framework that integrates image description, diagnostic step extraction, and automatic quality judgment into a coordinated pipeline, and we evaluate the reliability of the system through a human verification process. This framework enables efficient construction of fine-grained and clinically aligned supervision at scale. Building on the curated data, we propose In-Context Diffusion Transformer (IC-DiT), a layout-aware generative model that incorporates spatial layouts, textual descriptions, and visual embeddings into a unified diffusion transformer. Through hierarchical multimodal attention, IC-DiT maintains global semantic coherence while accurately preserving structural and morphological details. Extensive experiments on five histopathology datasets show that IC-DiT achieves higher fidelity, stronger spatial controllability, and better diagnostic consistency than existing methods. In addition, the generated images serve as effective data augmentation resources for downstream tasks such as cancer classification and survival analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。