提出推理时校准方法,提升时间分类的决策可靠性。
Inference-Time Decision Calibration for Temporal Classification

- 分离表示与校准,冻结主干网络,仅在推理时干预
- 多尺度残差分支在噪声环境中有显著增益,尤其短时序数据
- 针对互补证据设计动态校准,适合模型决策不充分场景
时间分类错误常被归因于表征不足,但也可能源于证据到决策的转换问题。本文提出表示-校准分解框架:保持训练好的主干分类器不变,引入两个推理时干预模块——保守的残差多尺度分支,增加辅助逻辑值;以及事后分支感知校准器,在决策时重新组合主干与残差证据。该设计可区分缺失的时间证据与未充分利用的决策级证据,无需重训练主干。在FI-2010、PTB-XL、UCI-HAR、MHEALTH和HARTH数据集上测试显示,性能提升具有强依赖性:残差多尺度证据在噪声大或表征受限场景中效果最佳,特别是短时序的FI-2010和较弱的循环主干;分支感知校准则在主干与辅助逻辑值含互补信息但未被原始决策规则充分利用时更有效。近饱和场景下两种干预收益有限。结果表明,时间分类不仅是表征学习问题,更是如何信任、融合与校准多视角证据的问题。
原文摘要 · Abstract (English)
Temporal classification errors are often treated as representation failures, but they can also arise from how available evidence is converted into decisions. This paper proposes a representation--calibration decomposition for temporal classification. We keep a trained native classifier frozen and separate two inference-time interventions: a conservative residual multi-scale branch that adds auxiliary logits to the native prediction, and a post-hoc branch-aware calibrator that recombines native and residual evidence at decision time. This design distinguishes missing temporal evidence from underused decision-level evidence without retraining the backbone. Across FI-2010, PTB-XL, UCI-HAR, MHEALTH, and HARTH, we find that gains are strongly regime-dependent. Residual multi-scale evidence is most useful in noisy or representation-limited settings, especially short-horizon FI-2010 and weaker recurrent backbones, while branch-aware calibration helps when native and auxiliary logits contain complementary evidence not fully exploited by the raw decision rule. Near-saturated settings show limited gains from either intervention. These results suggest that temporal classification should be understood not only as representation learning, but also as the problem of trusting, combining, and calibrating evidence from multiple views.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。