用不确定相似性学习,保护隐私还提升分类效果。
Learning from Uncertain Similarity and Unlabeled Data
- 引入不确定性成分缓解标签泄露风险
- 理论证明达到最优参数收敛速率
- 适合隐私敏感场景的弱监督学习
现有基于相似性的弱监督学习方法通常依赖数据对间精确的相似性标注,可能无意中暴露敏感标签信息,引发隐私风险。为缓解此问题,我们提出一种新框架USimUL,其中每个相似性对都嵌入不确定性成分以降低标签泄露。本文提出一个无偏风险估计器,可从不确定相似性和未标记数据中学习。此外,理论上证明该估计器能达到统计最优的参数收敛速率。在基准和真实世界数据集上的大量实验表明,相比传统方法,本方法在分类性能上表现更优。
原文摘要 · Abstract (English)
Existing similarity-based weakly supervised learning approaches often rely on precise similarity annotations between data pairs, which may inadvertently expose sensitive label information and raise privacy risks. To mitigate this issue, we propose Uncertain Similarity and Unlabeled Learning (USimUL), a novel framework where each similarity pair is embedded with an uncertainty component to reduce label leakage. In this paper, we propose an unbiased risk estimator that learns from uncertain similarity and unlabeled data. Additionally, we theoretically prove that the estimator achieves statistically optimal parametric convergence rates. Extensive experiments on both benchmark and real-world datasets show that our method achieves superior classification performance compared to conventional similarity-based approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。