比较多种异常检测方法在纵向数据中的表现。
An empirical comparison of some outlier detection methods with longitudinal data
- 对比传统统计与机器学习的异常检测方法
- 新方法更灵活,适用于多维数据
- 结果以分数形式输出,便于评估异常程度
本文研究纵向数据中的异常检测问题。通过将官方统计中常用的传统方法与数据挖掘和机器学习领域的基于距离或二分决策树的方法进行比较,应用于不同统计单元的面板调查数据。传统方法较为简单,可直接识别潜在异常点,但需满足特定假设;而新方法仅提供一个与异常可能性相关的得分。所有方法均需设置调参参数,但最新方法更具灵活性,且在某些情况下比传统方法更有效,同时可处理多维数据。
原文摘要 · Abstract (English)
This note investigates the problem of detecting outliers in longitudinal data. It compares well-known methods used in official statistics with proposals from the fields of data mining and machine learning that are based on the distance between observations or binary partitioning trees. This is achieved by applying the methods to panel survey data related to different types of statistical units. Traditional methods are quite simple, enabling the direct identification of potential outliers, but they require specific assumptions. In contrast, recent methods provide only a score whose magnitude is directly related to the likelihood of an outlier being present. All the methods require the user to set a number of tuning parameters. However, the most recent methods are more flexible and sometimes more effective than traditional methods. In addition, these methods can be applied to multidimensional data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。