arXiv:2510.19716cs.CV2025-10被引 2

从视频中提取稳定可解释的动态变量,抗干扰能力强。

LyTimeT: Towards Robust and Interpretable State-Variable Discovery

  • 用时空注意力自动编码器聚焦关键区域,抑制无关干扰
  • 通过李雅普诺夫正则化降低预测误差,提升长期稳定性
  • 适合需要物理可解释性的复杂系统建模任务

从高维视频中提取系统的真正动力学变量,因背景运动、遮挡和纹理变化等视觉干扰而困难。我们提出 LyTimeT,一种两阶段可解释变量提取框架,学习鲁棒且稳定的动态系统潜在表示。第一阶段,采用基于 TimeSformer 的时空自编码器,利用全局注意力聚焦动态相关区域,抑制干扰因素,实现抗干扰的潜在状态学习与准确的长时序视频预测。第二阶段,对学习到的潜在空间进行探测,通过线性相关性分析选取最符合物理意义的维度,并使用基于李雅普诺夫的稳定性正则化改进转移动态,强制收缩以减少滚动预测中的误差累积。在五个合成基准和四个真实世界动力系统(包括混沌现象)上的实验表明,LyTimeT 在互信息和内在维度估计上最接近真实值,在背景扰动下保持不变,并在 CNN 基础(TIDE)和纯变压器基线中取得最低分析均方误差。结果证明,结合时空注意力与稳定性约束,可得到既准确又物理可解释的预测模型。

原文摘要 · Abstract (English)

Extracting the true dynamical variables of a system from high-dimensional video is challenging due to distracting visual factors such as background motion, occlusions, and texture changes. We propose LyTimeT, a two-phase framework for interpretable variable extraction that learns robust and stable latent representations of dynamical systems. In Phase 1, LyTimeT employs a spatio-temporal TimeSformer-based autoencoder that uses global attention to focus on dynamically relevant regions while suppressing nuisance variation, enabling distraction-robust latent state learning and accurate long-horizon video prediction. In Phase 2, we probe the learned latent space, select the most physically meaningful dimensions using linear correlation analysis, and refine the transition dynamics with a Lyapunov-based stability regularizer to enforce contraction and reduce error accumulation during roll-outs. Experiments on five synthetic benchmarks and four real-world dynamical systems, including chaotic phenomena, show that LyTimeT achieves mutual information and intrinsic dimension estimates closest to ground truth, remains invariant under background perturbations, and delivers the lowest analytical mean squared error among CNN-based (TIDE) and transformer-only baselines. Our results demonstrate that combining spatio-temporal attention with stability constraints yields predictive models that are not only accurate but also physically interpretable.

动态系统可解释性时空注意力李雅普诺夫

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。