通过扰动对比样例,提升分类边界附近的学习效率
Learning Half-Spaces from Perturbed Contrastive Examples
- 用可变噪声函数控制对比样本的扰动程度
- 在特定条件下,学习速度比传统方法快得多
- 适合研究主动学习与边界敏感样本的场景
我们研究一种两步对比样例查询机制,其中每个被查询或采样的带标签样本都会配对一个相反标签的对比样本。与先前理想化设定不同,本文引入由非减噪声函数 $f$ 参数化的扰动机制,其中扰动大小由样本到决策边界的距离 $d$ 决定。直观上,靠近边界的样本获得更高质量的对比样本。我们在两种设定下分析该模型:(i) 最大扰动幅度固定;(ii) 扰动随机。针对一维阈值及有界域上均匀分布的半空间,我们刻画了主动与被动对比样本复杂度随函数 $f$ 的变化规律。结果显示,在 $f$ 满足一定条件时,对比样例的存在可显著降低渐近查询复杂度与期望查询复杂度。
原文摘要 · Abstract (English)
We study learning under a two-step contrastive example oracle, as introduced by Mansouri et. al. (2025), where each queried (or sampled) labeled example is paired with an additional contrastive example of opposite label. While Mansouri et al. assume an idealized setting, where the contrastive example is at minimum distance of the originally queried/sampled point, we introduce and analyze a mechanism, parameterized by a non-decreasing noise function $f$, under which this ideal contrastive example is perturbed. The amount of perturbation is controlled by $f(d)$, where $d$ is the distance of the queried/sampled point to the decision boundary. Intuitively, this results in higher-quality contrastive examples for points closer to the decision boundary. We study this model in two settings: (i) when the maximum perturbation magnitude is fixed, and (ii) when it is stochastic. For one-dimensional thresholds and for half-spaces under the uniform distribution on a bounded domain, we characterize active and passive contrastive sample complexity in dependence on the function $f$. We show that, under certain conditions on $f$, the presence of contrastive examples speeds up learning in terms of asymptotic query complexity and asymptotic expected query complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。