提出风格不变性评估模型预测可靠性,提升测试时自适应的可信度
Towards Reliable Test-Time Adaptation: Style Invariance as a Correctness Likelihood
- 通过变换输入风格并检测预测一致性来估计正确概率
- 相比传统方法平均降低13个百分点校准误差
- 无需反向传播,可适配任意测试时自适应方法
测试时自适应(TTA)能高效调整部署模型,但在高风险领域常导致预测不确定性校准不佳。现有校准方法多假设模型或分布固定,在真实动态测试环境下性能下降。本文提出风格不变性作为正确性似然(SICL),通过测量风格变换后预测的一致性,以实例级方式估计正确概率,仅需前向传播,为即插即用、无需反向传播的校准模块,兼容任意TTA方法。在四个基线、五种TTA方法及三种模型架构下,两种真实场景的综合评估表明,SICL相较传统校准方法平均降低13个百分点校准误差。
原文摘要 · Abstract (English)
Test-time adaptation (TTA) enables efficient adaptation of deployed models, yet it often leads to poorly calibrated predictive uncertainty - a critical issue in high-stakes domains such as autonomous driving, finance, and healthcare. Existing calibration methods typically assume fixed models or static distributions, resulting in degraded performance under real-world, dynamic test conditions. To address these challenges, we introduce Style Invariance as a Correctness Likelihood (SICL), a framework that leverages style-invariance for robust uncertainty estimation. SICL estimates instance-wise correctness likelihood by measuring prediction consistency across style-altered variants, requiring only the model's forward pass. This makes it a plug-and-play, backpropagation-free calibration module compatible with any TTA method. Comprehensive evaluations across four baselines, five TTA methods, and two realistic scenarios with three model architecture demonstrate that SICL reduces calibration error by an average of 13 percentage points compared to conventional calibration approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。