arXiv:2510.08938cs.LGcs.CV2025-10被引 1

动态调整不确定度校准参数,提升模型在变化数据下的可靠性。

Bi-level Meta-Policy Control for Dynamic Uncertainty Calibration in Evidential Deep Learning

  • 用双层优化动态调节损失函数中的关键系数
  • 在多个任务上显著改善预测准确率与不确定性校准
  • 适合高风险决策场景的自适应模型部署

传统证据深度学习方法依赖静态超参数进行不确定度校准,难以适应动态数据分布,导致高风险决策任务中校准效果差、泛化能力弱。为此,我们提出元策略控制器(MPC),一种动态元学习框架,可自适应调整KL散度系数和狄利克雷先验强度以实现最优不确定度建模。具体而言,内层循环通过动态配置的损失函数更新模型参数,外层循环则由策略网络基于多目标奖励(平衡预测精度与不确定度质量)优化KL系数及类别特定的狄利克雷先验强度。相较于固定先验的方法,本方法的可学习狄利克雷先验能灵活适配类别分布与训练动态。大量实验表明,MPC显著提升了各类任务下模型预测的可靠性与校准性,改进了不确定度校准效果、预测准确率,并在基于置信度剔除样本后仍保持更强性能保留能力。

原文摘要 · Abstract (English)

Traditional Evidence Deep Learning (EDL) methods rely on static hyperparameter for uncertainty calibration, limiting their adaptability in dynamic data distributions, which results in poor calibration and generalization in high-risk decision-making tasks. To address this limitation, we propose the Meta-Policy Controller (MPC), a dynamic meta-learning framework that adjusts the KL divergence coefficient and Dirichlet prior strengths for optimal uncertainty modeling. Specifically, MPC employs a bi-level optimization approach: in the inner loop, model parameters are updated through a dynamically configured loss function that adapts to the current training state; in the outer loop, a policy network optimizes the KL divergence coefficient and class-specific Dirichlet prior strengths based on multi-objective rewards balancing prediction accuracy and uncertainty quality. Unlike previous methods with fixed priors, our learnable Dirichlet prior enables flexible adaptation to class distributions and training dynamics. Extensive experimental results show that MPC significantly enhances the reliability and calibration of model predictions across various tasks, improving uncertainty calibration, prediction accuracy, and performance retention after confidence-based sample rejection.

不确定度校准元学习证据深度学习动态适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。