arXiv:2607.18279cs.LGcs.AI2026-07

用频谱特征增强时间序列分类的可靠性判断,避免盲目信任高置信度结果。

Beyond Output-Space Calibration: Spectral Evidence Bundling for Selective Reliability Estimation in Time-Series Classification

论文配图:Beyond Output-Space Calibration: Spectral Evidence Bundling for Selective Reliability Estimation in Time-Series Classification
图 1 · 摘自论文原文
  • 融合输出置信度与全样本频谱特征(如能量、熵、周期支持等)生成可靠性评分
  • 在8个数据集上将相关性AURC提升至0.786,虚假高置信错误率降至0.094
  • 通过验证门控机制确保频谱调整仅在有效时启用,保障安全性

时间序列分类的后处理校准通常仅重映射输出得分,但部署决策如信任、拒判和复审依赖于当前时序信号是否支撑置信预测。本文揭示三个可靠性缺口:相同置信值可能隐藏不同时序支持、平均校准会遗漏虚假高置信错误、输出空间校准难以追溯输入关联。提出一种验证门控的固定标签可靠性策略,在不改变主干预测的前提下,结合输出线索与全样本频谱描述符(包括带能量、熵、峰值主导性、周期支持、相位稳定性)生成标量可靠性估计与带级诊断证据。验证门仅在正确性排序提升且不违反[email protected]或AURC容忍度时启用频谱条件;否则回退至更安全的输出空间基线。在八个异构UCR/UEA数据集、八种时间序列骨干模型及标准校准器上,无约束方法在匹配评估子集上提升固定标签选择性可靠性指标,将Corr-AURC从0.693提高至0.779。验证门控策略进一步将Corr-AURC提升至0.786,并将[email protected]降低至0.094。结果表明,时间序列分类器的可靠性估计需融合输出置信度与频谱证据,而验证门控可防止无效的频谱调节。

原文摘要 · Abstract (English)

Post-hoc calibration for time-series classification usually remaps output scores, but deployment decisions such as trust, abstention, and review depend on whether a confident prediction is supported by the current temporal signal. We address three time-series reliability gaps: identical confidence values can hide different temporal support, average calibration can miss false high-confidence errors, and output-space recalibration offers limited input-linked auditability. We introduce a validation-gated fixed-label reliability policy that keeps the backbone prediction unchanged while estimating whether it should be trusted. The method combines output-side cues with whole-sample spectral descriptors, including band energy, entropy, peak dominance, period support, and phase stability, to form a scalar reliability estimate and diagnostic band-level evidence. A validation gate enables spectral conditioning only when correctness ranking improves without breaching [email protected] or AURC tolerances; otherwise it reverts to the safer output-space baseline. Across eight heterogeneous UCR/UEA datasets, eight time-series backbone families, and standard recalibrators, the unconstrained method improves fixed-label selective-reliability metrics on the matched evaluation subset, raising Corr-AURC from 0.693 to 0.779. The validation-gated policy further improves Corr-AURC to 0.786 and reduces [email protected] to 0.094. These results suggest that reliability estimation for time-series classifiers benefits from bundling output confidence with spectral evidence, while validation gating prevents unsupported spectral conditioning.

时间序列可靠性估计频谱分析校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。