显式分解视频中物体运动,提升预测质量
On the Benefits of Instance Decomposition in Video Prediction Models
- 将场景中的物体分离建模,而非整体统一处理
- 在合成与真实数据集上均实现更高质量预测
- 适合关注视频预测细节与可解释性的研究者
视频预测对机器人和自动驾驶车辆等智能体至关重要,使其能提前预判并应对突发情况。当前主流方法通常隐式联合建模场景动态,未显式分解为独立物体。这可能次优,因为动态场景中每个物体的运动模式通常相对独立。本文研究在潜在变换器视频预测模型中显式分解物体所带来的收益。我们在合成与真实数据集上进行细致且受控的实验,结果表明,对动态场景进行分解后,预测质量显著优于同等容量但无分解结构的模型。
原文摘要 · Abstract (English)
Video prediction is a crucial task for intelligent agents such as robots and autonomous vehicles, since it enables them to anticipate and act early on time-critical incidents. State-of-the-art video prediction methods typically model the dynamics of a scene jointly and implicitly, without any explicit decomposition into separate objects. This is challenging and potentially sub-optimal, as every object in a dynamic scene has their own pattern of movement, typically somewhat independent of others. In this paper, we investigate the benefit of explicitly modeling the objects in a dynamic scene separately within the context of latent-transformer video prediction models. We conduct detailed and carefully-controlled experiments on both synthetic and real-world datasets; our results show that decomposing a dynamic scene leads to higher quality predictions compared with models of a similar capacity that lack such decomposition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。