arXiv:2512.15505cs.CVeess.IV2025-12被引 3

重新评估脑影像配准零样本性能,发现深度学习并非通用优势。

The LUMirage: An independent evaluation of zero-shot performance in the LUMIR challenge

  • 独立验证零样本泛化能力,采用严谨评估协议
  • 对新对比度数据性能下降显著,效应量达0.7-1.5
  • 对预处理敏感且高分辨率下无法运行,适合临床实测者关注

LUMIR挑战是评估变形图像配准方法在大规模神经影像数据上表现的重要基准。尽管该挑战表明现代深度学习方法在T1加权MRI上达到可比精度,还宣称其对未见对比度和分辨率具有卓越的零样本泛化能力,这与深度学习领域关于域偏移的公认理解相悖。本文通过严格的评估协议,独立重新评估这些零样本主张,并排除潜在仪器偏差。结果揭示更复杂的图景:(1) 深度学习方法在分布内T1w图像上表现与迭代优化相当,甚至在人类近缘物种(猕猴)上亦表现优异,体现任务理解提升;(2) 在分布外对比度(T2、T2*、FLAIR)上性能显著下降,Cohen's d值为0.7–1.5,表明对下游临床工作流程有显著实际影响;(3) 深度方法在高分辨率数据上存在可扩展性限制,无法运行于0.6 mm等距图像,而迭代方法随分辨率提升表现更好;(4) 深度方法对预处理选择高度敏感。这些结果符合领域内关于域偏移的成熟文献,提示对零样本普适性声称需审慎对待。我们主张采用反映真实临床与研究工作流的评估协议,而非可能无意中偏向特定方法类别的条件。

原文摘要 · Abstract (English)

The LUMIR challenge represents an important benchmark for evaluating deformable image registration methods on large-scale neuroimaging data. While the challenge demonstrates that modern deep learning methods achieve competitive accuracy on T1-weighted MRI, it also claims exceptional zero-shot generalization to unseen contrasts and resolutions, assertions that contradict established understanding of domain shift in deep learning. In this paper, we perform an independent re-evaluation of these zero-shot claims using rigorous evaluation protocols while addressing potential sources of instrumentation bias. Our findings reveal a more nuanced picture: (1) deep learning methods perform comparably to iterative optimization on in-distribution T1w images and even on human-adjacent species (macaque), demonstrating improved task understanding; (2) however, performance degrades significantly on out-of-distribution contrasts (T2, T2*, FLAIR), with Cohen's d scores ranging from 0.7-1.5, indicating substantial practical impact on downstream clinical workflows; (3) deep learning methods face scalability limitations on high-resolution data, failing to run on 0.6 mm isotropic images, while iterative methods benefit from increased resolution; and (4) deep methods exhibit high sensitivity to preprocessing choices. These results align with the well-established literature on domain shift and suggest that claims of universal zero-shot superiority require careful scrutiny. We advocate for evaluation protocols that reflect practical clinical and research workflows rather than conditions that may inadvertently favor particular method classes.

图像配准零样本神经影像域偏移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。