arXiv:2606.00630cs.CVstat.ML2026-06被引 2

用术中超声生成类似MRI图像,提升手术导航精度。

A Systematic Benchmark of Intraoperative Ultrasound-to-MR Synthesis for Brain Tumour Surgery

论文配图:A Systematic Benchmark of Intraoperative Ultrasound-to-MR Synthesis for Brain Tumour Surgery
图 1 · 摘自论文原文
  • 构建统一基准测试6种生成模型在4种推理模式下的表现。
  • 发现感知质量与下游分割效果相关性最高,传统指标易误导。
  • 扩散模型在2.5D模式下分割精度最优,适合实际手术场景。

术中超声(ioUS)是脑肿瘤手术中一种灵活且经济的影像方式,但其图像解读困难:成像平面非标准、存在特定伪影,且外观与术前MRI差异显著,而术前MRI正是手术规划、分割模型和外科医生经验依赖的基础。从ioUS合成类似MRI的图像,可使现有MRI基础设施在术中复用,无需额外扫描。以往研究多孤立评估单一架构,尚无涵盖不同架构范式、推理模式和下游任务目标的统一基准。本文在公开的ReMIND数据集上开展系统性评测(76例患者;153对ioUS/T2w和104对ioUS/FLAIR配对数据;60/16患者级训练/保留划分)。六种生成器(四种GAN基线:Pix2Pix、SwinPix2Pix、CycleGAN、CUT;基于Transformer的ResViT;以及少步扩散模型SynDiff)在四种推理模式(2D、2.5D、2D+3D精炼、全3D)和两种目标(仅T2w;T2w+FLAIR多任务)下共执行48组实验。除图像保真度指标(SSIM、PSNR、MAE、LPIPS)外,还采用nnU-Net v2进行下游分割评估(肿瘤和切除腔),并按组织学分级和再手术情况进行子组分析。结果表明:无任一架构在所有维度占优;关键发现是感知质量(LPIPS)与下游实用性最相关(r=-0.66, p<0.001),而高SSIM反而与较差实用性相关(r=-0.64, p<0.001);SynDiff-2.5D在下游分割中表现最佳(U_Dice=0.55)。因此,应优先报告感知质量和下游任务性能,而非仅依赖全局SSIM;架构选择应根据手术阶段、患者病史和临床目标调整。

原文摘要 · Abstract (English)

Intraoperative ultrasound (ioUS) is a versatile, cost-effective modality in brain tumour surgery, but its interpretation is difficult: acquisition planes are non-standard, artefacts are modality-specific, and its appearance differs markedly from the preoperative MRI on which surgical-planning tools, segmentation models and the surgeon's experience rely. Synthesising MRI-like images from ioUS could let this MRI-based infrastructure be reused intraoperatively without an extra scan. Most prior work evaluates a single architecture in isolation; to our knowledge, no benchmark has spanned architectural paradigms, inference regimes and downstream-task endpoints under a common protocol. We address this gap on the public ReMIND data set (76 patients; 153 paired ioUS/T2w and 104 paired ioUS/FLAIR studies; 60/16 patient-level train/held-out split). Six generators (four GAN baselines: Pix2Pix, SwinPix2Pix, CycleGAN, CUT; the transformer-augmented ResViT; and the few-step diffusion model SynDiff) were each trained under four inference regimes (2D, 2.5D, 2D + 3D-refinement, full-3D) and two targets (T2w only; T2w + FLAIR multi-task), yielding 48 experiments. Image-fidelity metrics (SSIM, PSNR, MAE, LPIPS) were complemented by an nnU-Net v2 downstream segmentation evaluation (tumour and resection cavity) and by subgroup analyses by histological grade and reoperation. No architecture dominated every axis, and, critically, perceptual quality tracked downstream utility most closely (LPIPS, r=-0.66, p<0.001), whereas higher SSIM was associated with worse utility (r=-0.64, p<0.001); SynDiff-2.5D best preserved downstream segmentation (U_Dice=0.55). Perceptual and downstream-task metrics should therefore be reported alongside or in preference to global SSIM, and architecture choice conditioned on surgical phase, patient history and clinical objective.

医学图像生成超声合成手术导航扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。