arXiv:2510.23261cs.LG2025-10NeurIPS被引 1

提出两种可解释的时序分段评估方法,更精准识别错误类型与位置。

Toward Interpretable Evaluation Measures for Time Series Segmentation

  • 引入加权调整兰德指数(WARI)和状态匹配得分(SMS),考虑误差位置与类型。
  • 在真实与合成数据上验证,能准确评估分段质量并揭示错误来源。
  • 适合需要理解模型错误机制的科研人员和工程应用者。

时序分段是分析跨领域时间数据(如人体活动识别、能源监控)的基础任务。尽管已有大量先进方法,其性能评估仍严重受限。现有指标多关注变点准确性或依赖点对点度量(如调整兰德指数ARI),无法刻画分段质量,忽略错误性质,且可解释性差。本文提出两种新评估指标:WARI(加权调整兰德指数),考虑分段误差的位置;SMS(状态匹配得分),细粒度识别并评分四种基本类型的分段错误,并支持按错误类型加权。在合成与真实世界基准上实证验证表明,两者不仅更准确评估分段质量,还能揭示传统方法无法获取的错误来源与类型信息。

原文摘要 · Abstract (English)

Time series segmentation is a fundamental task in analyzing temporal data across various domains, from human activity recognition to energy monitoring. While numerous state-of-the-art methods have been developed to tackle this problem, the evaluation of their performance remains critically limited. Existing measures predominantly focus on change point accuracy or rely on point-based measures such as Adjusted Rand Index (ARI), which fail to capture the quality of the detected segments, ignore the nature of errors, and offer limited interpretability. In this paper, we address these shortcomings by introducing two novel evaluation measures: WARI (Weighted Adjusted Rand Index), that accounts for the position of segmentation errors, and SMS (State Matching Score), a fine-grained measure that identifies and scores four fundamental types of segmentation errors while allowing error-specific weighting. We empirically validate WARI and SMS on synthetic and real-world benchmarks, showing that they not only provide a more accurate assessment of segmentation quality but also uncover insights, such as error provenance and type, that are inaccessible with traditional measures.

时序分段评估指标可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。