多尺度距离度量比对时间序列更准,不依赖点对点匹配。
Beyond Point Matching: Evaluating Multiscale Dubuc Distance for Time Series Similarity
- 基于多尺度分析,跳过传统点对点对齐。
- 在95个数据集上,多数场景下优于DTW,提升显著。
- 适合处理复杂时间序列的分类任务,实用性强。
时间序列是高维复杂的数据对象,其高效搜索与索引一直是数据挖掘中的长期挑战。本文基于近期提出的相似性度量方法——多尺度杜布克距离(Multiscale Dubuc Distance, MDD),系统评估其相对于广泛使用的动态时间规整(Dynamic Time Warping, DTW)的优势与局限。MDD 的核心创新在于:跨多个时间尺度评估时间序列相似性,并避免点对点对齐。我们通过模拟实验及来自 UCR 时间序列分类基准的 95 个数据集验证假设,结果表明,在诸多场景中 MDD 显著优于 DTW,且对其性能瓶颈有深入解析。进一步将两种方法应用于一项具有挑战性的实际分类任务,结果显示 MDD 显著超越 DTW,凸显其实际应用价值。
原文摘要 · Abstract (English)
Time series are high-dimensional and complex data objects, making their efficient search and indexing a longstanding challenge in data mining. Building on a recently introduced similarity measure, namely Multiscale Dubuc Distance (MDD), this paper investigates its comparative strengths and limitations relative to the widely used Dynamic Time Warping (DTW). MDD is novel in two key ways: it evaluates time series similarity across multiple temporal scales and avoids point-to-point alignment. We demonstrate that in many scenarios where MDD outperforms DTW, the gains are substantial, and we provide a detailed analysis of the specific performance gaps it addresses. We provide simulations, in addition to the 95 datasets from the UCR archive, to test our hypotheses. Finally, we apply both methods to a challenging real-world classification task and show that MDD yields a significant improvement over DTW, underscoring its practical utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。