arXiv:2501.16718cs.LG2025-01ICLR被引 5

用哈密顿蒙特卡洛生成高质量异常数据,提升模型对分布外样本的检测能力。

Outlier Synthesis via Hamiltonian Monte Carlo for Out-of-Distribution Detection

  • 基于马尔可夫链采样,在分布内数据上合成多样异常样本
  • 采样接受率接近1,效率高且生成样本质量好
  • 无需真实异常数据,适合缺乏标注异常样本的场景

分布外(OOD)检测对于构建可信可靠的机器学习系统至关重要。近期通过引入辅助异常数据训练的方法在提升检测能力方面表现优异,但这些方法严重依赖大量高质量的真实异常数据。部分先前方法尝试通过合成虚拟异常来缓解该问题,但因采样策略单调或生成模型参数量大而面临生成质量差或成本高的困境。本文提出哈密顿蒙特卡洛异常合成框架(HamOS),将合成过程视为马尔可夫链采样。仅基于分布内数据,该框架可在特征空间中广泛遍历,生成多样且具代表性的异常样本,使模型暴露于多种潜在分布外情形。同时,具有接近1的采样接受率,显著提升效率。在标准与大规模基准上的实证结果表明,所提HamOS在效果与效率上均优于现有最先进方法。

原文摘要 · Abstract (English)

Out-of-distribution (OOD) detection is crucial for developing trustworthy and reliable machine learning systems. Recent advances in training with auxiliary OOD data demonstrate efficacy in enhancing detection capabilities. Nonetheless, these methods heavily rely on acquiring a large pool of high-quality natural outliers. Some prior methods try to alleviate this problem by synthesizing virtual outliers but suffer from either poor quality or high cost due to the monotonous sampling strategy and the heavy-parameterized generative models. In this paper, we overcome all these problems by proposing the Hamiltonian Monte Carlo Outlier Synthesis (HamOS) framework, which views the synthesis process as sampling from Markov chains. Based solely on the in-distribution data, the Markov chains can extensively traverse the feature space and generate diverse and representative outliers, hence exposing the model to miscellaneous potential OOD scenarios. The Hamiltonian Monte Carlo with sampling acceptance rate almost close to 1 also makes our framework enjoy great efficiency. By empirically competing with SOTA baselines on both standard and large-scale benchmarks, we verify the efficacy and efficiency of our proposed HamOS.

OOD检测异常生成哈密顿采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。