arXiv:2602.13066cs.CV2026-02中稿 · ISBI 2026

提出可检测生成模型泄露训练数据的精准指标,适用于医疗影像安全审查。

A Calibrated Memorization Index (MI) for Detecting Training Data Leakage in Generative MRI Models

  • 基于MRI基础模型提取特征,融合多层相似性计算
  • 在三个数据集上实现样本级近似完美漏检率
  • 适合关注医学图像隐私与模型安全的研究者

生成式图像模型会复制训练数据中的图像,用于医学图像生成时可能引发隐私问题。本文提出一种校准的逐样本度量方法,用于检测记忆与重复现象。该方法利用MRI基础模型提取图像特征,聚合多层白化后的最近邻相似性,并映射为有界得分:过拟合/新颖性指数(ONI)和记忆指数(MI)。在包含可控重复率及典型图像增强的三个MRI数据集上,该指标能稳健检测重复内容,且跨数据集结果更一致;在样本层面,实现近乎完美的重复检测能力。

原文摘要 · Abstract (English)

Image generative models are known to duplicate images from the training data as part of their outputs, which can lead to privacy concerns when used for medical image generation. We propose a calibrated per-sample metric for detecting memorization and duplication of training data. Our metric uses image features extracted using an MRI foundation model, aggregates multi-layer whitened nearest-neighbor similarities, and maps them to a bounded \emph{Overfit/Novelty Index} (ONI) and \emph{Memorization Index} (MI) scores. Across three MRI datasets with controlled duplication percentages and typical image augmentations, our metric robustly detects duplication and provides more consistent metric values across datasets. At the sample level, our metric achieves near-perfect detection of duplicates.

生成模型医学影像隐私安全记忆检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。