arXiv:2607.25984cs.CVcs.LG2026-07中稿 · ECCV

用概率分布建模场景未来运动,实现快速精准的多轨迹预测。

Schrödinger's Cat: Probabilistic Representation and Prediction of Potential Scene Kinematics

论文配图:Schrödinger's Cat: Probabilistic Representation and Prediction of Potential Scene Kinematics
图 1 · 摘自论文原文
  • 学习场景未来运动的概率潜空间表示,支持多轨迹联合采样。
  • 采样速度比现有模型快97倍,密度估计快100倍。
  • 适合需要实时交互与不确定性分析的机器人规划任务。

从部分观测中预测场景演化,需推理多种可能的未来而非单一轨迹。现有方法或生成以外观为主的视频,或仅采样少量轨迹而未显式建模运动分布。我们提出目标感知的未来运动潜分布表示(GARFIELD),一种基于图像和可选稀疏时空约束的结构化时空潜表示,用于建模可能未来的分布。该表示支持所有轨迹的联合采样,并通过高效的确定性密度解码器直接访问底层运动分布。因此,未来运动的不确定性可定位到特定场景元素与时间点,并随额外约束逐步细化。实验表明,该方法在运动规划性能上媲美大型视频生成模型,同时采样速度提升97倍;运动密度估计速度比蒙特卡洛采样快两个数量级,支持交互式探索与不确定性感知规划。

原文摘要 · Abstract (English)

Predicting how a scene may evolve from partial observations requires reasoning about multiple possible futures rather than committing to a single trajectory. Existing approaches either generate appearance-dominated video predictions or sample a small number of trajectories without explicitly modeling the distribution of possible motion. We introduce Goal-Aware Representations of Future kInEmatic Latent Distributions (GARFIELD), a probabilistic model of scene kinematics that learns a structured spatio-temporal latent representation of the distribution over possible futures given an image and optional spatio-temporally sparse constraints. The same latent representation enables both joint sampling of all trajectories and direct access to the underlying motion distribution through an efficient deterministic density decoder. As a result, uncertainty about future motion can be localized to specific scene elements and timesteps and progressively refined through additional constraints. Experiments demonstrate strong motion planning performance competitive with large video generation models while sampling trajectories $97\times$ faster. Our method further estimates motion densities two orders of magnitude faster than Monte-Carlo sampling from motion generation models, enabling interactive exploration and uncertainty-aware planning.

运动预测概率建模机器人规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。