用物理规律做先验,让模型学会可解释的视频表示并生成新场景。
Interpretable Representation Learning from Videos using Nonlinear Priors
- 用非线性噪声模型替代传统高斯先验,引入物理规律约束。
- 在摆锤、弹簧质量等真实物理视频上成功学习出正确变量。
- 可干预变量生成符合物理的新视频,适合需要可解释性的研究者。
学习视觉数据的可解释表示是使机器决策对人类可理解并提升训练分布外泛化能力的关键挑战。本文提出一种深度学习框架,允许用户为视频指定非线性先验(如牛顿力学),使模型学习可解释的潜在变量,并用于生成训练时未见过的假设场景视频。通过将变分自编码器(VAE)的先验从简单的各向同性高斯扩展为任意非线性时序加性噪声模型(ANM),该方法可描述多种过程(如牛顿物理)。我们提出一种新颖的线性化方法,构建高斯混合模型(GMM)近似先验,并推导出后验与先验GMM间KL散度的数值稳定蒙特卡洛估计。在多个真实世界物理视频数据集上验证:包括摆锤、弹簧质量系统、下落物体和脉冲星(旋转中子星)。对每项实验设定相应物理先验,结果显示模型能正确学习对应物理变量。模型训练完成后,可通过干预改变如振幅或空气阻力等物理变量,生成符合物理规律的未见场景视频。
原文摘要 · Abstract (English)
Learning interpretable representations of visual data is an important challenge, to make machines' decisions understandable to humans and to improve generalisation outside of the training distribution. To this end, we propose a deep learning framework where one can specify nonlinear priors for videos (e.g. of Newtonian physics) that allow the model to learn interpretable latent variables and use these to generate videos of hypothetical scenarios not observed at training time. We do this by extending the Variational Auto-Encoder (VAE) prior from a simple isotropic Gaussian to an arbitrary nonlinear temporal Additive Noise Model (ANM), which can describe a large number of processes (e.g. Newtonian physics). We propose a novel linearization method that constructs a Gaussian Mixture Model (GMM) approximating the prior, and derive a numerically stable Monte Carlo estimate of the KL divergence between the posterior and prior GMMs. We validate the method on different real-world physics videos including a pendulum, a mass on a spring, a falling object and a pulsar (rotating neutron star). We specify a physical prior for each experiment and show that the correct variables are learned. Once a model is trained, we intervene on it to change different physical variables (such as oscillation amplitude or adding air drag) to generate physically correct videos of hypothetical scenarios that were not observed previously.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。