arXiv:2608.13418stat.MLcs.LG2026-08

通过最大化样本分布与污染分布的Wasserstein距离,筛选出异常样本以恢复干净数据分布。

Wasserstein Filtering: A Sample Selection Method for Robust Distribution Learning

论文配图:Wasserstein Filtering: A Sample Selection Method for Robust Distribution Learning
图 1 · 摘自论文原文
  • 基于Wasserstein距离选择几何上影响大的离群点进行剔除
  • 在合成数据和扩散模型生成任务中,重污染下仍能有效提升生成质量
  • 无需假设模型结构,适用于各类下游任务的鲁棒预处理

当数据集部分样本被污染时,目标是恢复原始干净分布。本文提出Wasserstein Filtering(WF)框架,通过剔除可疑样本并利用剩余样本的经验分布估计目标分布。核心思想是选取一个子集,使其经验分布与完全污染的经验分布之间的Wasserstein距离最大,从而优先分离并移除几何上具有影响力的离群点。为使优化可计算,提出三种算法:边际筛选方案SinkMarg,以及两种联合优化算法SinkWF和SlicedWF,分别利用熵正则最优传输和切片Wasserstein近似。理论上,引入了远端排除与局部投影(FELP)污染模型,刻画由分离离群点和局部不可区分扰动组成的污染。在此模型下,证明了WF估计器在协方差有界的分布族上达到极小极大最优。在合成数据、基准异常检测套件及扩散模型的鲁棒生成学习中大量实验表明,WF是一种高度实用、模型无关的预处理工具,在重污染下表现出色,兼具良好离群点检测性能与显著的下游收益。

原文摘要 · Abstract (English)

Given a dataset where a portion of the samples are contaminated, our goal is to recover the underlying clean population distribution. To this end, we propose Wasserstein Filtering (WF), a novel sample selection framework that discards a fraction of suspicious samples and estimates the target distribution using the empirical measure of the remaining data. The core insight is to select a subset of samples whose empirical distribution maximizes its Wasserstein distance to the fully contaminated empirical distribution, thereby preferentially isolating and removing geometrically influential outliers. To render this optimization computationally tractable, we introduce three algorithms: a marginal screening scheme, SinkMarg, and two joint optimization algorithms, SinkWF and SlicedWF, leveraging entropic optimal transport and sliced Wasserstein approximations, respectively. On the theoretical front, we introduce the Far Exclusion and Local Projection (FELP) contamination model, which characterizes corruptions consisting of well-separated outliers and locally indistinguishable perturbations. Under this model, we prove that the WF estimator achieves minimax optimality over distribution families with bounded covariance. Extensive numerical experiments on synthetic datasets, benchmark anomaly detection suites, and robust generative learning with diffusion models demonstrate that WF serves as a highly practical, model-agnostic preprocessing tool. It delivers competitive outlier detection performance and provides substantial downstream benefits for generative modeling under heavy contamination.

分布学习异常检测生成模型最优传输

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。