arXiv:2508.20656cs.LG2025-08

用符号动力学和组合数据增强,让临床时间序列生成更真实合成数据。

Compositionality in Time Series: A Proof of Concept using Symbolic Dynamics and Compositional Data Augmentation

  • 将临床时间序列看作潜在生理状态的有序序列,挖掘其生成规律。
  • 合成数据训练模型性能接近真实数据,且优于随机增强方法。
  • 适合临床数据稀缺场景,对器官衰竭评分预测提升明显。

本文研究自然现象的时间序列是否由有序且规律的潜在状态序列生成。聚焦临床时间序列,探讨临床指标能否被解释为有意义生理状态的有序演化。揭示其组合结构后,可生成合成数据以缓解临床时间序列预测中数据稀疏、资源匮乏的问题,并深化对临床数据的理解。我们从数据生成过程出发定义时间序列的组合性,随后研究数据驱动方法以重构基础状态与组合规则。通过两个源自领域自适应视角的实证测试评估方法有效性:分别比较在原始与合成数据上训练/测试的时序预测模型的期望风险相似性。实验表明,基于组合合成数据训练的模型在测试集表现与真实数据相当;在合成测试数据上的评估结果也与真实测试数据相似,优于基于随机化的数据增强。下游任务中,完全基于组合合成数据训练的模型在序列器官衰竭评估(SOFA)分数预测任务上显著优于基于原始数据训练的模型。

原文摘要 · Abstract (English)

This work investigates whether time series of natural phenomena can be understood as being generated by sequences of latent states which are ordered in systematic and regular ways. We focus on clinical time series and ask whether clinical measurements can be interpreted as being generated by meaningful physiological states whose succession follows systematic principles. Uncovering the underlying compositional structure will allow us to create synthetic data to alleviate the notorious problem of sparse and low-resource data settings in clinical time series forecasting, and deepen our understanding of clinical data. We start by conceptualizing compositionality for time series as a property of the data generation process, and then study data-driven procedures that can reconstruct the elementary states and composition rules of this process. We evaluate the success of this methods using two empirical tests originating from a domain adaptation perspective. Both tests infer the similarity of the original time series distribution and the synthetic time series distribution from the similarity of expected risk of time series forecasting models trained and tested on original and synthesized data in specific ways. Our experimental results show that the test set performance achieved by training on compositionally synthesized data is comparable to training on original clinical time series data, and that evaluation of models on compositionally synthesized test data shows similar results to evaluating on original test data, outperforming randomization-based data augmentation. An additional downstream evaluation of the prediction task of sequential organ failure assessment (SOFA) scores shows significant performance gains when model training is entirely based on compositionally synthesized data compared to training on original data.

时间序列合成数据临床预测组合性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。