无需位姿先验,用前馈4D高斯点云预测自动驾驶未来场景。
Envision4D: Envisioning Visual Futures via Feed-forward 4D Gaussian Splatting for Autonomous Driving

- 通过迭代去噪预测未来相机位姿,实现无监督位姿推断。
- 引入时序注意力与运动提升机制,有效捕捉非线性动态变化。
- 适合需要高精度未来视图合成的自动驾驶场景研究者。
预测动态场景的未来演化对自动驾驶至关重要。然而,现有前馈范式主要针对插值设计,扩展到未来外推时,在大位移下易产生鬼影伪影,且受限于简化运动假设或严格未来先验。为此,我们提出Envision4D,一种完全自监督的无位姿前馈未来外推框架。具体地,我们引入未来位姿预测模块,通过迭代去噪过程推断未来相机参数。为捕捉非线性动力学,提出层内时序注意力,并采用条件运动提升,将高度不确定的外推过程转化为稳健的关系映射。最后,通过渐进式训练策略稳定无监督运动学习,抑制误差累积。大量实验表明,Envision4D在未来视图合成上达到领先性能,显著优于现有方法。
原文摘要 · Abstract (English)
Forecasting the future evolution of dynamic scenes is crucial in autonomous driving. However, existing feed-forward paradigms are primarily designed for interpolation. When extended to future extrapolation, they suffer from ghosting artifacts under large displacements and are constrained by simplified motion assumptions or strict future priors. To overcome these challenges, we propose Envision4D, a fully self-supervised feed-forward framework for pose-free future extrapolation. Specifically, we introduce a Future Pose Prediction module that infers future camera parameters via an iterative denoising process. Furthermore, to capture non-linear dynamics, we propose In-layer Temporal Attention and employ Conditioned Motion Lifting, which transforms the highly uncertain extrapolation process into robust relational mappings. Finally, a Progressive Training Strategy is utilized to stabilize unsupervised motion learning against error accumulation. Extensive experiments demonstrate that Envision4D achieves state-of-the-art performance, significantly outperforming existing methods in future view synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。