用森林相似度生成时间序列嵌入,发现误分类点与异常点关联更强。
Forest Proximities for Time Series
- 基于森林相似度构建时间序列嵌入,替代传统距离度量
- 误分类样本与异常点的关联性显著强于近邻分类器
- 适合时间序列分类、异常检测与可解释性研究
RF-GAP 是一种改进的随机森林相似度度量。本文提出 PF-GAP,将其扩展至邻近森林(proximity forests),一种准确高效的时序分类模型。我们结合多维缩放(Multi-Dimensional Scaling)使用森林相似度,生成单变量时序的向量嵌入,并与多种时序距离度量所得嵌入进行对比。同时,将森林相似度与局部离群因子(Local Outlier Factor)结合,探究误分类点与异常点之间的关系,对比基于时序距离的近邻分类器。结果表明,森林相似度表现出比近邻分类器更强的误分类点与异常点关联性。
原文摘要 · Abstract (English)
RF-GAP has recently been introduced as an improved random forest proximity measure. In this paper, we present PF-GAP, an extension of RF-GAP proximities to proximity forests, an accurate and efficient time series classification model. We use the forest proximities in connection with Multi-Dimensional Scaling to obtain vector embeddings of univariate time series, comparing the embeddings to those obtained using various time series distance measures. We also use the forest proximities alongside Local Outlier Factors to investigate the connection between misclassified points and outliers, comparing with nearest neighbor classifiers which use time series distance measures. We show that the forest proximities seem to exhibit a stronger connection between misclassified points and outliers than nearest neighbor classifiers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。