arXiv:2504.09085cs.LG2025-04

提出真实场景下基于噪声标签数据的超参数优化框架

crowd-hpo: Realistic Hyperparameter Optimization and Benchmarking for Learning from Crowds with Noisy Labels

  • 仅用噪声众包标签构建验证集进行超参数筛选
  • 提升模型在真实测试集上的泛化性能
  • 适合评估众包学习方法的研究者使用

众包标注是获取类别标签的低成本方案,但标签存在噪声。现有方法通常使用默认超参数配置或依赖带真实标签的验证集进行调优,前者导致性能不公,后者不现实。为此,我们提出 crowd-hpo 框架,支持仅利用噪声众包标签的验证数据,选择表现优异的超参数配置。大量实验表明,该框架选出的配置能显著提升神经网络在独立测试集(含真实标签)上的泛化能力。因此,在研究中引入此类筛选标准,对实现更公平、更真实的众包学习基准评估至关重要。

原文摘要 · Abstract (English)

Crowdworking is a cost-efficient solution for acquiring class labels. Since these labels are subject to noise, various approaches to learning from crowds have been proposed. Typically, these approaches are evaluated with default hyperparameter configurations, resulting in unfair and suboptimal performance, or with hyperparameter configurations tuned via a validation set with ground truth class labels, representing an often unrealistic scenario. Moreover, both setups can produce different approach rankings, complicating study comparisons. Therefore, we introduce crowd-hpo as a framework for evaluating approaches to learning from crowds in combination with criteria to select well-performing hyperparameter configurations with access only to noisy crowd-labeled validation data. Extensive experiments with neural networks demonstrate that these criteria select hyperparameter configurations, which improve the learning from crowd approaches' generalization performances, measured on separate test sets with ground truth labels. Hence, incorporating such criteria into experimental studies is essential for enabling fairer and more realistic benchmarking.

众包学习超参数优化噪声标签基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。