arXiv:2508.07723cs.CV2025-08被引 1

用三元组连接重加权提升小样本生成数据质量,有效抑制噪声图像。

Enhancing Small-Scale Dataset Expansion with Triplet-Connection-based Sample Re-Weighting

  • 基于三元组连接设计重加权机制,自动识别并降低噪声生成样本权重。
  • 在6个自然图像数据集上平均提升7.9%,3个医学数据集上平均提升3.4%。
  • 可兼容任意生成数据增强方法,适合医疗等小样本视觉任务应用。

在医疗诊断等实际应用中,计算机视觉模型性能常受限于可用图像的稀缺性。利用预训练生成模型扩充数据集是有效方案,但生成过程不可控且自然语言描述模糊,可能导致噪声图像生成。重加权可通过为噪声图像分配低权重来缓解此问题。本文首次从理论上分析了三类生成图像的监督方式,并据此提出TriReWeight——一种基于三元组连接的样本重加权方法,用于增强生成数据增强效果。理论上,该方法可与任意生成数据增强方法结合,且不会降低其性能;其泛化误差逼近最优阶 $O(ig rac{\\(d\ln n)}{n}\\sqrt{}$。实验验证了理论分析正确性,表明该方法在六个自然图像数据集上平均优于现有最先进方法7.9%,在三个医学数据集上平均提升3.4%。同时验证了其对不同生成数据增强方法均有性能提升作用。

原文摘要 · Abstract (English)

The performance of computer vision models in certain real-world applications, such as medical diagnosis, is often limited by the scarcity of available images. Expanding datasets using pre-trained generative models is an effective solution. However, due to the uncontrollable generation process and the ambiguity of natural language, noisy images may be generated. Re-weighting is an effective way to address this issue by assigning low weights to such noisy images. We first theoretically analyze three types of supervision for the generated images. Based on the theoretical analysis, we develop TriReWeight, a triplet-connection-based sample re-weighting method to enhance generative data augmentation. Theoretically, TriReWeight can be integrated with any generative data augmentation methods and never downgrade their performance. Moreover, its generalization approaches the optimal in the order $O(\sqrt{d\ln (n)/n})$. Our experiments validate the correctness of the theoretical analysis and demonstrate that our method outperforms the existing SOTA methods by $7.9\%$ on average over six natural image datasets and by $3.4\%$ on average over three medical datasets. We also experimentally validate that our method can enhance the performance of different generative data augmentation methods.

数据增强生成模型医学图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。