arXiv:2512.13164cs.CVcs.AI2025-12

用文本生成病理图像,解决数据少、质量差难题

A Semantically Enhanced Generative Foundation Model Improves Pathological Image Synthesis

  • 基于280万图文对训练,通过关联约束机制抑制语义漂移
  • 生成30种癌症病理图像,经病理科医生验证质量达标
  • 可结合分割图精准控制组织结构,适合罕见癌种研究

病理科人工智能的发展受限于多样且高质量标注数据集的匮乏。生成模型虽具潜力,但存在语义不稳定和形态幻觉问题,影响诊断可靠性。为此,我们提出首个针对病理学的文本到图像生成基础模型CRAFTS,采用双阶段训练策略,在约280万张图像-描述对上训练,引入新型对齐机制以抑制语义漂移,确保生物真实性。该模型生成涵盖30种癌症类型的多样化病理图像,其质量通过客观指标和病理科医生评估双重验证。此外,使用CRAFTS增强的数据集显著提升多种临床任务性能,包括分类、跨模态检索、自监督学习和视觉问答。进一步地,将CRAFTS与ControlNet结合,可基于核分割图、荧光图像等输入精确控制组织架构。通过克服数据稀缺与隐私问题,CRAFTS提供了无限量、多样化的标注组织学数据,有效推动稀有及复杂癌变表型诊断工具的构建。

原文摘要 · Abstract (English)

The development of clinical-grade artificial intelligence in pathology is limited by the scarcity of diverse, high-quality annotated datasets. Generative models offer a potential solution but suffer from semantic instability and morphological hallucinations that compromise diagnostic reliability. To address this challenge, we introduce a Correlation-Regulated Alignment Framework for Tissue Synthesis (CRAFTS), the first generative foundation model for pathology-specific text-to-image synthesis. By leveraging a dual-stage training strategy on approximately 2.8 million image-caption pairs, CRAFTS incorporates a novel alignment mechanism that suppresses semantic drift to ensure biological accuracy. This model generates diverse pathological images spanning 30 cancer types, with quality rigorously validated by objective metrics and pathologist evaluations. Furthermore, CRAFTS-augmented datasets enhance the performance across various clinical tasks, including classification, cross-modal retrieval, self-supervised learning, and visual question answering. In addition, coupling CRAFTS with ControlNet enables precise control over tissue architecture from inputs such as nuclear segmentation masks and fluorescence images. By overcoming the critical barriers of data scarcity and privacy concerns, CRAFTS provides a limitless source of diverse, annotated histology data, effectively unlocking the creation of robust diagnostic tools for rare and complex cancer phenotypes.

病理图像生成生成模型医学数据增强文本到图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。