利用标签引导的树模型近邻关系,提升时间序列分类中的缺失值填补效果。
Label-Guided Imputation via Forest-Based Proximities for Improved Time Series Classification
- 基于标签信息构建树模型近邻,指导缺失值填补。
- 相比传统方法,分类准确率显著提升,即使填补值不完全真实。
- 适合有类别标签的时间序列数据缺失处理任务。
时间序列数据中缺失值问题普遍存在。现有大多数填补方法忽视了时间序列对应的类别标签信息。本文提出一种面向时间序列分类任务的缺失值填补框架,其中每个时间序列均关联一个类别标签。通过利用高精度的监督学习模型,提取其树结构生成的近邻度量,实现基于标签的条件性填补。实验表明,该方法虽填补值与真实值存在差异,但能提供更丰富的信息,显著提升分类准确率。
原文摘要 · Abstract (English)
Missing data is a common problem in time series data. Most methods for imputation ignore label information pertaining to the time series even if that information exists. In this paper, we provide a framework for missing data imputation in the context of time series classification, where each time series is associated with a categorical label. We define a means of imputing missing values conditional upon labels, the method being guided by powerful, existing supervised models designed for high accuracy in this task. From each model, we extract a tree-based proximity measure from which imputation can be applied. We show that imputation using this method generally provides richer information leading to higher classification accuracies, despite the imputed values differing from the true values.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。