arXiv:2603.01448cs.DBcs.LG2026-03被引 12

用深度网络提升数据序列相似性搜索精度,尤其在噪声大时表现更优。

SEAnet: A Deep Learning Architecture for Data Series Similarity Search

  • 设计新型深度嵌入近似方法,学习序列压缩特征。
  • 在7个真实与合成数据集上,相似性搜索准确率显著优于现有方法。
  • 适合处理高噪声、弱相关等复杂数据场景的科研与工业应用。

大规模数据序列分析中的核心操作是相似性搜索。现有研究表明,基于SAX的索引在相似性搜索任务中表现最佳,但在高频、弱相关、过度噪声或其他特定数据集特性下性能下降。本文提出一种基于深度神经网络的新颖数据序列摘要技术——深度嵌入近似(DEA)。同时,我们设计了专为学习DEA而生的SEAnet架构,首次将平方和保持性质引入深度网络设计。进一步结合SEAtrans编码器增强表达能力,并提出SEAsam与SEAsamE两种新采样策略,使SEAnet可在海量数据上高效训练。在7个不同来源的真实与合成数据集上的全面实验表明,使用SEAnet学习的DEA能提供高质量的数据序列摘要与相似性搜索结果。

原文摘要 · Abstract (English)

A key operation for massive data series collection analysis is similarity search. According to recent studies, SAX-based indexes offer state-of-the-art performance for similarity search tasks. However, their performance lags under high-frequency, weakly correlated, excessively noisy, or other dataset-specific properties. In this work, we propose Deep Embedding Approximation (DEA), a novel family of data series summarization techniques based on deep neural networks. Moreover, we describe SEAnet, a novel architecture especially designed for learning DEA, that introduces the Sum of Squares preservation property into the deep network design. We further enhance SEAnet with SEAtrans encoder. Finally, we propose novel sampling strategies, SEAsam and SEAsamE, that allow SEAnet to effectively train on massive datasets. Comprehensive experiments on 7 diverse synthetic and real datasets verify the advantages of DEA learned using SEAnet in providing high-quality data series summarizations and similarity search results.

序列搜索深度学习数据摘要嵌入学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。