找出高维时间序列中差异显著的变量与时间段。
Variable Selection for Comparing High-dimensional Time-Series Data
- 分段比较两组时间序列,筛选出差异显著的变量与区间。
- 在流体与交通模拟器对比中验证了方法的有效性。
- 适合模拟器验证、模型对比与参数敏感性分析。
针对长度和维度相同的两组多变量时间序列,提出一种变量与时间区间选择方法,以识别两者存在显著差异的区域。该方法将整个时间区间划分为多个子区间,在每个子区间内对两组样本进行分布比较,筛选出能区分分布的变量,并执行两样本检验。通过合成数据实验评估了该方法的有效性与局限性。在粒子流体模拟器的应用中,对比了深度神经网络模型与模拟器输出;在微观交通模拟器应用中,分析了改变模拟器参数对交通流的影响,验证了方法的实际价值。
原文摘要 · Abstract (English)
Given a pair of multivariate time-series data of the same length and dimensions, an approach is proposed to select variables and time intervals where the two series are significantly different. In applications where one time series is an output from a computationally expensive simulator, the approach may be used for validating the simulator against real data, for comparing the outputs of two simulators, and for validating a machine learning-based emulator against the simulator. With the proposed approach, the entire time interval is split into multiple subintervals, and on each subinterval, the two sample sets are compared to select variables that distinguish their distributions and a two-sample test is performed. The validity and limitations of the proposed approach are investigated in synthetic data experiments. Its usefulness is demonstrated in an application with a particle-based fluid simulator, where a deep neural network model is compared against the simulator, and in an application with a microscopic traffic simulator, where the effects of changing the simulator's parameters on traffic flows are analysed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。