arXiv:2412.00613cs.LGstat.ML2024-12被引 5

用全部数据学习表示,提升非参数两样本检验效果

A Unified Data Representation Learning for Non-parametric Two-sample Testing

  • 全数据自监督学习潜在表示,不依赖样本索引
  • 结合潜在表示与判别表示,显著提升检验功效
  • 适合需要高精度检验的机器学习研究者

有效数据表示的学习在非参数两样本检验中至关重要。传统方法将数据分为训练集和测试集,仅在训练集上学习表示,但近期理论研究表明,只要不使用样本索引,整个数据集均可用于表示学习,同时保证第一类错误控制。受此启发,我们提出表示学习两样本检验(RL-TST)框架:先对全数据进行纯自监督表示学习,捕获反映底层数据流形的内在表示(IRs);再基于这些表示训练判别模型,学习判别表示(DRs),从而融合结构信息与判别能力。大量实验表明,RL-TST通过利用测试集的数据流形信息并借助训练集寻找最优判别表示,显著优于现有代表性方法。

原文摘要 · Abstract (English)

Learning effective data representations has been crucial in non-parametric two-sample testing. Common approaches will first split data into training and test sets and then learn data representations purely on the training set. However, recent theoretical studies have shown that, as long as the sample indexes are not used during the learning process, the whole data can be used to learn data representations, meanwhile ensuring control of Type-I errors. The above fact motivates us to use the test set (but without sample indexes) to facilitate the data representation learning in the testing. To this end, we propose a representation-learning two-sample testing (RL-TST) framework. RL-TST first performs purely self-supervised representation learning on the entire dataset to capture inherent representations (IRs) that reflect the underlying data manifold. A discriminative model is then trained on these IRs to learn discriminative representations (DRs), enabling the framework to leverage both the rich structural information from IRs and the discriminative power of DRs. Extensive experiments demonstrate that RL-TST outperforms representative approaches by simultaneously using data manifold information in the test set and enhancing test power via finding the DRs with the training set.

两样本检验表示学习自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。