arXiv:2502.16736cs.LGcs.AI2025-02被引 1

根据引导信号的不确定性动态调整信任程度,提升模型在噪声数据下的学习效果。

Adaptive Conformal Guidance for Learning under Uncertainty

  • 用分割保形预测量化引导信号的不确定性,动态调节其影响权重。
  • 在网格世界导航中,奖励比最优基线高出6倍以上,收敛速度更快。
  • 适用于知识蒸馏、半监督学习等场景,适合处理不完美引导信号的任务。

引导学习在多种机器学习系统中表现有效,如监督学习中的标注数据、半监督学习中的伪标签、强化学习中的专家示范策略。然而,由于领域偏移和数据有限,引导信号可能包含噪声且泛化能力差。盲目信任存在噪声、不完整或与目标域不匹配的信号会导致性能下降。为此,我们提出自适应保形引导(AdaConG),一种简单而有效的方法,通过分割保形预测(CP)量化引导信号的不确定性,并据此动态调节其影响。通过自适应地应对引导信号的不确定性,AdaConG使模型减少对潜在误导性信号的依赖,从而提升学习性能。我们在知识蒸馏、半监督图像分类、网格世界导航和自动驾驶等多个任务上验证了该方法。实验结果表明,AdaConG在不完美引导下显著提升性能与鲁棒性,例如在网格世界导航中,其收敛速度更快,奖励超过最佳基线6倍以上。这些结果表明AdaConG是一种广泛适用的学习不确定性解决方案。

原文摘要 · Abstract (English)

Learning with guidance has proven effective across a wide range of machine learning systems. Guidance may, for example, come from annotated datasets in supervised learning, pseudo-labels in semi-supervised learning, and expert demonstration policies in reinforcement learning. However, guidance signals can be noisy due to domain shifts and limited data availability and may not generalize well. Blindly trusting such signals when they are noisy, incomplete, or misaligned with the target domain can lead to degraded performance. To address these challenges, we propose Adaptive Conformal Guidance (AdaConG), a simple yet effective approach that dynamically modulates the influence of guidance signals based on their associated uncertainty, quantified via split conformal prediction (CP). By adaptively adjusting to guidance uncertainty, AdaConG enables models to reduce reliance on potentially misleading signals and enhance learning performance. We validate AdaConG across diverse tasks, including knowledge distillation, semi-supervised image classification, gridworld navigation, and autonomous driving. Experimental results demonstrate that AdaConG improves performance and robustness under imperfect guidance, e.g., in gridworld navigation, it accelerates convergence and achieves over $6\times$ higher rewards than the best-performing baseline. These results highlight AdaConG as a broadly applicable solution for learning under uncertainty.

不确定性学习引导学习保形预测强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。