arXiv:2411.00842cs.CVcs.LG2024-11被引 1

用得分模型预测视频下一帧,能自动选择最可能的未来路径。

Video prediction using score-based conditional density estimation

  • 基于得分的隐式概率框架,从历史帧预测下一帧分布。
  • 在合成数据中成功处理遮挡边界,避免模糊平均,选择高概率轨迹。
  • 自然视频训练后能按可靠性加权预测证据,符合统计推断原理。

时间预测本质上具有不确定性,但在自然图像序列中表示这种模糊性是一个极具挑战性的高维概率推断问题。对于自然场景,维度灾难使得显式密度估计在统计和计算上都不可行。本文提出一种基于隐式回归的框架,用于学习和采样给定前序帧时下一帧的条件密度。我们证明,通过简单抗噪目标函数训练的序列到图像深度网络,能提取出适用于时间预测的自适应表示。合成实验表明,该得分框架可有效处理遮挡边界:与传统方法对分支时间轨迹进行平均不同,它能在多个可能轨迹中择优,更频繁地选择高概率选项。此外,对自然图像序列训练的网络分析显示,其表示会自动根据预测证据的可靠性进行加权,这是统计推断的核心特征。

原文摘要 · Abstract (English)

Temporal prediction is inherently uncertain, but representing the ambiguity in natural image sequences is a challenging high-dimensional probabilistic inference problem. For natural scenes, the curse of dimensionality renders explicit density estimation statistically and computationally intractable. Here, we describe an implicit regression-based framework for learning and sampling the conditional density of the next frame in a video given previous observed frames. We show that sequence-to-image deep networks trained on a simple resilience-to-noise objective function extract adaptive representations for temporal prediction. Synthetic experiments demonstrate that this score-based framework can handle occlusion boundaries: unlike classical methods that average over bifurcating temporal trajectories, it chooses among likely trajectories, selecting more probable options with higher frequency. Furthermore, analysis of networks trained on natural image sequences reveals that the representation automatically weights predictive evidence by its reliability, which is a hallmark of statistical inference

视频预测得分模型概率建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。