arXiv:2412.18256cs.LGcs.AI2024-12被引 20

解决开放环境下标签不一致数据对半监督学习的负面影响

Robust Semi-Supervised Learning in Open Environments

  • 提出在标签、特征、分布不一致时仍有效的半监督学习方法
  • 发现使用不一致无标签数据可能导致性能劣于有监督基线
  • 适合研究开放环境下的鲁棒机器学习的学者参考

半监督学习(SSL)旨在标签稀缺时利用无标签数据提升性能。传统研究多假设封闭环境,即有标签与无标签数据在标签、特征、分布等方面保持一致。然而,更实际的任务常处于开放环境,此时关键因素存在不一致。已有研究表明,使用不一致的无标签数据会导致严重性能下降,甚至劣于简单有监督基线。手动验证无标签数据质量不可行,因此研究开放环境下具有鲁棒性的半监督学习至关重要。本文简要综述该方向最新进展,聚焦标签、特征和数据分布不一致的处理技术,并提供评估基准。同时讨论了开放研究问题,供参考。

原文摘要 · Abstract (English)

Semi-supervised learning (SSL) aims to improve performance by exploiting unlabeled data when labels are scarce. Conventional SSL studies typically assume close environments where important factors (e.g., label, feature, distribution) between labeled and unlabeled data are consistent. However, more practical tasks involve open environments where important factors between labeled and unlabeled data are inconsistent. It has been reported that exploiting inconsistent unlabeled data causes severe performance degradation, even worse than the simple supervised learning baseline. Manually verifying the quality of unlabeled data is not desirable, therefore, it is important to study robust SSL with inconsistent unlabeled data in open environments. This paper briefly introduces some advances in this line of research, focusing on techniques concerning label, feature, and data distribution inconsistency in SSL, and presents the evaluation benchmarks. Open research problems are also discussed for reference purposes.

半监督学习开放环境鲁棒性数据不一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。