arXiv:2512.21058cs.CV2025-12中稿 · CVPR被引 3

用诊断语义令牌和原型控制生成病理图像,实现精准语义操控。

Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Control

  • 引入多流控制机制,融合文本、语义和原型信息生成图像。
  • 在68K高质量数据集上达到80.9的Patho-FID(领先第二名51%)。
  • 适合需要精细病理图像生成的研究者与医学人工智能开发者。

在计算病理学中,理解与生成能力发展脱节:先进理解模型已具备诊断级性能,而生成模型仍以像素模拟为主。进展受限于三大因素:缺乏大规模高质量图文语料;缺乏细粒度语义控制,迫使依赖非语义线索;术语异构性导致同一诊断概念表达多样,影响文本条件生成可靠性。我们提出UniPath,一种以语义驱动的病理图像生成框架,利用成熟的诊断理解能力实现可控生成。UniPath采用多流控制:原始文本流;高层语义流,通过可学习查询冻结的病理多模态大模型,提炼抗重述的诊断语义令牌,并将提示扩展为诊断感知属性包;原型流则通过原型库实现组件级形态控制。数据方面,我们构建了265万张图像-文本语料库及6.8万张精标注高质量子集,缓解数据稀缺问题。为全面评估,建立四层评估体系。大量实验表明,UniPath表现达到当前最优,包括80.9的Patho-FID(比第二好高出51%),细粒度语义控制接近真实图像的98.7%。数据与代码见https://github.com/Hanminghao/UniPath。

原文摘要 · Abstract (English)

In computational pathology, understanding and generation have evolved along disparate paths: advanced understanding models already exhibit diagnostic-level competence, whereas generative models largely simulate pixels. Progress remains hindered by three coupled factors: the scarcity of large, high-quality image-text corpora; the lack of precise, fine-grained semantic control, which forces reliance on non-semantic cues; and terminological heterogeneity, where diverse phrasings for the same diagnostic concept impede reliable text conditioning. We introduce UniPath, a semantics-driven pathology image generation framework that leverages mature diagnostic understanding to enable controllable generation. UniPath implements Multi-Stream Control: a Raw-Text stream; a High-Level Semantics stream that uses learnable queries to a frozen pathology MLLM to distill paraphrase-robust Diagnostic Semantic Tokens and to expand prompts into diagnosis-aware attribute bundles; and a Prototype stream that affords component-level morphological control via a prototype bank. On the data front, we curate a 2.65M image-text corpus and a finely annotated, high-quality 68K subset to alleviate data scarcity. For a comprehensive assessment, we establish a four-tier evaluation hierarchy tailored to pathology. Extensive experiments demonstrate UniPath's SOTA performance, including a Patho-FID of 80.9 (51% better than the second-best) and fine-grained semantic control achieving 98.7% of the real-image. The dataset and code can be obtained from https://github.com/Hanminghao/UniPath.

病理图像语义生成原型控制多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。