arXiv:2608.07176cs.CVcs.AI2026-08中稿 · the DCA-MI Worksho…

用内镜图像自学习对齐潜空间,生成更真实、结构更稳定的内镜影像。

Representation-driven Endoscopic Visual Embedding Alignment for Latent Generation

论文配图:Representation-driven Endoscopic Visual Embedding Alignment for Latent Generation
图 1 · 摘自论文原文
  • 基于内镜分布预训练编码器,直接对齐扩散模型潜空间与临床特征。
  • 在500万帧数据上训练,生成图像保真度高,结构一致性强。
  • 适合开发智能胃肠诊疗系统,尤其适用于图像合成与分割任务。

构建内镜领域的基础生成模型受限于自然图像与临床图像之间的差异,以及大模型训练的高计算成本。尽管表示对齐在通用计算机视觉中提升了效率,但在高度专业的内镜图像领域作用尚不明确。本文提出REVEAL(Representation-driven Endoscopic Visual Embedding Alignment),目前规模最大的内镜生成基础模型,基于包含500万帧的多中心数据集GastroNet-5M(GN-5M)训练。REVEAL不依赖外部域先验,而是直接在内镜分布上预训练编码器,将扩散潜空间与领域特异性视觉特征对齐,有效保留细微纹理与复杂解剖结构。除图像生成外,REVEAL还作为强大特征提取器,在多个基准测试中表现媲美甚至超越专为分类任务优化的EndoViT和Endo-FM等内镜基础模型,且在真实成像退化条件下展现强鲁棒性。该模型能生成高保真图像,并在修复与扩展等潜空间编辑任务中保持结构连贯性。其高容量主干降低了专用临床工具的计算门槛,为未来智能胃肠系统中的条件合成、分割及分布外检测提供开放、通用的基础支持。

原文摘要 · Abstract (English)

Developing foundation generative models for endoscopy is limited by the gap between natural and clinical images and the computational cost of training large Diffusion Transformers. Although representation alignment has improved efficiency in general computer vision, its role within the highly specialized endoscopic image space remains unclear. We introduce REVEAL (Representation-driven Endoscopic Visual Embedding Alignment), the largest generative foundation model for endoscopy to date, trained on GastroNet-5M (GN-5M), a multicenter dataset of 5 million endoscopic frames. Instead of depending on out-of-domain priors, REVEAL employs encoders pretrained directly on the endoscopic distribution to align diffusion latents with domain-specific visual features, preserving fine textures and intricate anatomical structures. Beyond image generation, REVEAL also serves as a powerful feature extractor; in multiple benchmarks, it delivers performance that is competitive with, and in several cases exceeds, endoscopic foundation models such as EndoViT and Endo-FM, specifically tuned for classification tasks, while demonstrating strong representation robustness under realistic imaging corruptions. REVEAL produces high-fidelity images and maintains robust structural coherence in latent-space edits such as inpainting and outpainting. This high-capacity backbone lowers the computational threshold for building specialized clinical tools, offering an open, versatile foundation for conditional synthesis, segmentation, and out-of-distribution detection in future intelligent gastroenterology systems.

内镜生成扩散模型医学图像特征对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。