arXiv:2510.01349cs.LGstat.ML2025-10被引 1

提出新方法诊断数据对称性偏差,揭示增强效果受限的深层原因。

To Augment or Not to Augment? Diagnosing Distributional Symmetry Breaking

  • 用双样本分类器检测原始数据与随机增强数据分布差异
  • 发现点云数据存在严重对称性偏差,影响模型泛化能力
  • 提醒研究者:对称性方法效果依赖数据偏差特性,需谨慎评估

对称性感知的机器学习方法(如数据增强、等变架构)假设变换后的数据在测试分布中具有高概率,从而提升泛化与样本效率。本文提出一种度量数据集对称性破坏程度的方法,通过双样本分类器区分原始数据与随机增强数据。在合成数据上验证后,发现多个基准点云数据集存在显著对称性偏差,构成严重数据偏见。理论证明,在无限特征极限下,分布对称性破坏即使标签真正不变,也会导致不变方法无法最优。实证显示,等变方法的效果因数据而异:当对称性偏差与标签相关时,其优势消失。这表明理解等变性需重新审视数据中的对称性偏见。

原文摘要 · Abstract (English)

Symmetry-aware methods for machine learning, such as data augmentation and equivariant architectures, encourage correct model behavior on all transformations (e.g. rotations or permutations) of the original dataset. These methods can improve generalization and sample efficiency, under the assumption that the transformed datapoints are highly probable, or "important", under the test distribution. In this work, we develop a method for critically evaluating this assumption. In particular, we propose a metric to quantify the amount of symmetry breaking in a dataset, via a two-sample classifier test that distinguishes between the original dataset and its randomly augmented equivalent. We validate our metric on synthetic datasets, and then use it to uncover surprisingly high degrees of symmetry-breaking in several benchmark point cloud datasets, constituting a severe form of dataset bias. We show theoretically that distributional symmetry-breaking can prevent invariant methods from performing optimally even when the underlying labels are truly invariant, for invariant ridge regression in the infinite feature limit. Empirically, the implication for symmetry-aware methods is dataset-dependent: equivariant methods still impart benefits on some symmetry-biased datasets, but not others, particularly when the symmetry bias is predictive of the labels. Overall, these findings suggest that understanding equivariance -- both when it works, and why -- may require rethinking symmetry biases in the data.

数据偏差对称性点云等变

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。