arXiv:2412.20617cs.LG2024-12

将时间序列转为字母序列,用生物信息学方法分析复杂数据模式。

Converting Time Series Data to Numeric Representations Using Alphabetic Mapping and k-mer strategy

  • 用26个字母分区间映射时间序列值,生成字符序列。
  • 在真实数据上实现序列分类,验证了方法有效性。
  • 适合想用生物序列技术分析时间序列的研究者。

在数据分析与生物信息学领域,将时间序列数据以类似生物序列的方式表示,为应用序列分析技术提供了新思路。通过将时间序列信号转换为类分子序列的表示形式,可借助生物信息学中成熟的基于k-mer的分析方法,提升对复杂非线性时间序列中隐藏模式和关系的识别能力。本文提出一种独特的字母映射方法:将时间序列数值划分为26个区间,分别对应英文字母表中的26个字母,每个数值据此映射为特定字符。该转换使传统生物信息学中的序列分析算法能直接应用于时间序列数据。我们通过将真实世界的时间序列信号转化为字符序列,并进行序列分类任务,验证了该方法的有效性。所得序列可用于多种基于序列的分析技术,为时间序列数据的表示与分析提供了新视角。

原文摘要 · Abstract (English)

In the realm of data analysis and bioinformatics, representing time series data in a manner akin to biological sequences offers a novel approach to leverage sequence analysis techniques. Transforming time series signals into molecular sequence-type representations allows us to enhance pattern recognition by applying sophisticated sequence analysis techniques (e.g. $k$-mers based representation) developed in bioinformatics, uncovering hidden patterns and relationships in complex, non-linear time series data. This paper proposes a method to transform time series signals into biological/molecular sequence-type representations using a unique alphabetic mapping technique. By generating 26 ranges corresponding to the 26 letters of the English alphabet, each value within the time series is mapped to a specific character based on its range. This conversion facilitates the application of sequence analysis algorithms, typically used in bioinformatics, to analyze time series data. We demonstrate the effectiveness of this approach by converting real-world time series signals into character sequences and performing sequence classification. The resulting sequences can be utilized for various sequence-based analysis techniques, offering a new perspective on time series data representation and analysis.

时间序列序列分析生物信息学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。