arXiv:2501.00910cs.LGcs.AI2025-01AAAI被引 17

让生成的时间序列更真实,既保个体质量也保整体分布。

Population Aware Diffusion for Time Series Generation

  • 设计新训练方法,显式保留数据整体统计特性。
  • 生成数据的维度间相关性分布误差降低5.9倍。
  • 适合需要真实统计数据的下游预测任务使用。

扩散模型在生成高质量时间序列数据方面表现优异。然而,现有方法多关注个体数据的真实性,忽视了整个数据集的群体层面属性,如各维度的取值分布及不同维度间的功能依赖关系(如交叉相关性,CC)分布。例如,在生成房屋能耗时间序列时,应保持室外温度与厨房温度的分布及其相互间的交叉相关性分布。保留这些群体属性对维护数据集的统计洞察力、减少模型偏差以及增强下游时间序列预测任务至关重要,但常被现有模型忽略。因此,现有模型生成的数据往往存在分布偏移。本文提出一种新的时间序列生成模型——人口感知扩散模型(PaD-TS),能更好保留群体属性。其核心创新包括:1)一种显式纳入时间序列群体属性保护的新型训练方法;2)一种能更好捕捉数据结构的双通道编码器架构。在主流基准数据集上的实验证明,PaD-TS可使真实数据与合成数据间的平均交叉相关性分布偏移分数降低5.9倍,同时在个体真实性上达到与顶尖模型相当的水平。

原文摘要 · Abstract (English)

Diffusion models have shown promising ability in generating high-quality time series (TS) data. Despite the initial success, existing works mostly focus on the authenticity of data at the individual level, but pay less attention to preserving the population-level properties on the entire dataset. Such population-level properties include value distributions for each dimension and distributions of certain functional dependencies (e.g., cross-correlation, CC) between different dimensions. For instance, when generating house energy consumption TS data, the value distributions of the outside temperature and the kitchen temperature should be preserved, as well as the distribution of CC between them. Preserving such TS population-level properties is critical in maintaining the statistical insights of the datasets, mitigating model bias, and augmenting downstream tasks like TS prediction. Yet, it is often overlooked by existing models. Hence, data generated by existing models often bear distribution shifts from the original data. We propose Population-aware Diffusion for Time Series (PaD-TS), a new TS generation model that better preserves the population-level properties. The key novelties of PaD-TS include 1) a new training method explicitly incorporating TS population-level property preservation, and 2) a new dual-channel encoder model architecture that better captures the TS data structure. Empirical results in major benchmark datasets show that PaD-TS can improve the average CC distribution shift score between real and synthetic data by 5.9x while maintaining a performance comparable to state-of-the-art models on individual-level authenticity.

时间序列生成扩散模型分布保持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。