arXiv:2608.13729cs.CV2026-08

合成数据在小众医学领域效果有限,反而可能误导模型。

Limitations of Synthetic Data Generation in Specialized Data-Scarce Domains

论文配图:Limitations of Synthetic Data Generation in Specialized Data-Scarce Domains
图 1 · 摘自论文原文
  • 用扩散模型生成图像来补数据,对比非生成方法
  • 五项创伤分类任务中,合成数据未超越传统增强方法
  • 生成图像常失真或过度简化,适合特定任务但不真实

基于扩散模型的生成技术推动了合成图像在视觉任务中的应用,以缓解数据稀缺问题。尽管在ImageNet等自然图像基准上表现良好,但在高差异性、数据稀疏的真实场景中效果仍不明确。本文聚焦于与常见数据集差异大且新增数据成本高的领域,评估两类生成式数据扩展方法——分布建模与样本扰动——在五个创伤分类任务上的表现(采用按受试者划分的训练-验证集)。结果表明,无一生成方法在所有任务中持续优于强基线的非生成数据增强方法。特征空间分析揭示了常见失败模式:记忆或崩溃、分布漂移,以及生成看似合理但结构简化的典型实例,这些实例比真实数据更易分类。

原文摘要 · Abstract (English)

Advances in diffusion-based generative models have motivated the use of synthetic image generation to alleviate data scarcity in vision tasks. While this strategy has shown promise in natural image benchmarks such as ImageNet, its effectiveness in sparse, high-variance real-world domains remains unclear. In this work, we focus on domains where images differ substantially from common image datasets and additional data are expensive to obtain. Against non-generative data augmentation baselines, we evaluate the downstream classifier performance improvements yielded by two schools of generative sparse data extension: distribution modeling and sample perturbation. Across five trauma classification tasks using subject-wise train--validation splits, no generative approach consistently outperforms a strong non-generative baseline. Feature-space analysis reveals recurring failure modes: memorization or collapse, distributional drift, and generation of visually plausible but simplified canonical instances that are easier to classify than real data.

合成数据医学图像扩散模型数据稀缺

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。