arXiv:2606.07498cs.CV2026-06

用权重扰动生成对比样本,避免科学数据因图像变换失真。

Implicit Data Synthesis for Contrastive Unsupervised Data Augmentation

论文配图:Implicit Data Synthesis for Contrastive Unsupervised Data Augmentation
图 1 · 摘自论文原文
  • 不扰动数据本身,改扰动网络权重生成对比样本
  • 在流星雷达数据上提升对比学习性能
  • 适合对数据扰动敏感的科学观测场景

科学观测产生大量未标注数据,人工标注成本高,无监督学习因而重要。对比学习可从无标签数据中提取结构化表征。传统方法依赖数据空间增强生成合成样本,但对科学观测数据而言,此类扰动可能根本性改变数据结构。本文提出通过扰动网络权重而非原始数据来生成对比样本,更有效保留数据本质结构。基于SimCLR框架,在流星雷达观测数据上验证该方法,实验显示在相同评估协议下性能提升。

原文摘要 · Abstract (English)

Scientific observations generate large quantities of unlabeled data which is laborious to hand-label, making unsupervised learning techniques valuable for processing datasets. Among these approaches, contrastive learning provides a convenient mechanism for extracting structural representations from unannotated datasets. For natural imagery, the general approach is to use a variety of data-space augmentation methods in order to generate synthetic samples; however, for scientific observations data-space perturbations can fundamentally alter the underlying data. Our proposed method is to generate contrastive samples by perturbing the network weights rather than the underlying data, thus more closely preserving the structure of the data. We demonstrate this technique using a SimCLR-based pipeline applied over radar observations of meteors, and show performance gains under matched protocols.

对比学习无监督学习科学数据数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。