发现视频生成模型只学了30.8%真实数据,影响下游任务性能。
Uncovering Hidden Subspaces in Video Diffusion Models Using Re-Identification
- 用重识别模型分析生成视频的潜在空间覆盖度。
- 实测仅30.8%训练视频被模型有效学习。
- 适合关注生成数据质量与医疗影像隐私的科研人员。
潜在空间中的视频扩散模型虽能生成高质量、时序一致的视频,吸引大量关注,但其在医疗等敏感领域应用时仍存隐私隐患。尽管合成数据可促进安全共享,但基于合成数据训练的下游模型性能仍逊于真实数据。本文指出,这可能源于采样空间仅为训练视频的子空间,实际有效数据量减少。此外,长视频生成时序一致性下降也是因素之一。我们首次证明在潜在空间训练隐私保护模型更高效且泛化更强。进一步提出使用重识别模型(此前用于隐私过滤)分析生成数据覆盖范围,仅需在生成器潜在空间训练即可。该方法可用于评估生成模型的忠实度。以心脏超声为例,结果表明潜在视频扩散模型仅学习到最多30.8%的真实训练视频,解释了下游任务性能不足的原因。
原文摘要 · Abstract (English)
Latent Video Diffusion Models can easily deceive casual observers and domain experts alike thanks to the produced image quality and temporal consistency. Beyond entertainment, this creates opportunities around safe data sharing of fully synthetic datasets, which are crucial in healthcare, as well as other domains relying on sensitive personal information. However, privacy concerns with this approach have not fully been addressed yet, and models trained on synthetic data for specific downstream tasks still perform worse than those trained on real data. This discrepancy may be partly due to the sampling space being a subspace of the training videos, effectively reducing the training data size for downstream models. Additionally, the reduced temporal consistency when generating long videos could be a contributing factor. In this paper, we first show that training privacy-preserving models in latent space is computationally more efficient and generalize better. Furthermore, to investigate downstream degradation factors, we propose to use a re-identification model, previously employed as a privacy preservation filter. We demonstrate that it is sufficient to train this model on the latent space of the video generator. Subsequently, we use these models to evaluate the subspace covered by synthetic video datasets and thus introduce a new way to measure the faithfulness of generative machine learning models. We focus on a specific application in healthcare echocardiography to illustrate the effectiveness of our novel methods. Our findings indicate that only up to 30.8% of the training videos are learned in latent video diffusion models, which could explain the lack of performance when training downstream tasks on synthetic data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。