arXiv:2509.04819eess.IVcs.CV2025-09

AURAD可生成高保真胸部X光片与伪语义掩码,提升医学图像合成的可控性与临床实用性。

AURAD: Anatomy-Pathology Unified Radiology Synthesis with Progressive Representations

  • 分阶段生成伪掩码并引导图像合成,结合解剖结构与病理特征
  • 78%合成图像被放射科医生判定为真实,40%分割结果具临床价值
  • 适合医学影像数据增强、病灶检测与分割等下游任务

医学图像合成已成为数据稀缺临床场景中扩充数据集、提升模型泛化能力的重要策略。然而,由于高质量标注有限及跨数据集域偏移,细粒度且可控的合成仍具挑战。现有方法多针对自然图像或明确肿瘤设计,难以泛化至胸部X光片——其疾病模式形态多样,且紧密交织于解剖结构之中。为此,我们提出AURAD,一种可控的放射学合成框架,能联合生成高保真胸部X光片与伪语义掩码。不同于依赖随机采样掩码的方法(限制多样性、可控性与临床相关性),本方法学习生成捕捉多病灶共存与解剖-病理一致性的掩码。采用渐进式流程:先根据临床提示生成基于解剖结构的伪掩码,再用于指导图像合成;同时利用预训练医疗专家模型过滤输出,确保临床合理性。除视觉真实性外,合成掩码还可作为检测与分割等下游任务的标签,弥合生成建模与真实临床应用间的鸿沟。大量实验与盲评放射科医生评估表明,该方法在多种任务与数据集上均具有效性与泛化能力:78%合成图像被认证为真实,超过40%预测的分割重叠被评定为具有临床价值。所有代码、预训练模型及合成数据集将在发表后公开。

原文摘要 · Abstract (English)

Medical image synthesis has become an essential strategy for augmenting datasets and improving model generalization in data-scarce clinical settings. However, fine-grained and controllable synthesis remains difficult due to limited high-quality annotations and domain shifts across datasets. Existing methods, often designed for natural images or well-defined tumors, struggle to generalize to chest radiographs, where disease patterns are morphologically diverse and tightly intertwined with anatomical structures. To address these challenges, we propose AURAD, a controllable radiology synthesis framework that jointly generates high-fidelity chest X-rays and pseudo semantic masks. Unlike prior approaches that rely on randomly sampled masks-limiting diversity, controllability, and clinical relevance-our method learns to generate masks that capture multi-pathology coexistence and anatomical-pathological consistency. It follows a progressive pipeline: pseudo masks are first generated from clinical prompts conditioned on anatomical structures, and then used to guide image synthesis. We also leverage pretrained expert medical models to filter outputs and ensure clinical plausibility. Beyond visual realism, the synthesized masks also serve as labels for downstream tasks such as detection and segmentation, bridging the gap between generative modeling and real-world clinical applications. Extensive experiments and blinded radiologist evaluations demonstrate the effectiveness and generalizability of our method across tasks and datasets. In particular, 78% of our synthesized images are classified as authentic by board-certified radiologists, and over 40% of predicted segmentation overlays are rated as clinically useful. All code, pre-trained models, and the synthesized dataset will be released upon publication.

医学图像合成胸部X光可控生成伪掩码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。