arXiv:2608.09493cs.CV2026-08中稿 · ECCV被引 1

通过视图感知的混合推理,提升交通视频未来帧预测的几何稳定性。

GeoRoute: Geometry-Aware Hybrid Inference for Traffic Future-Frame Prediction

论文配图:GeoRoute: Geometry-Aware Hybrid Inference for Traffic Future-Frame Prediction
图 1 · 摘自论文原文
  • 利用多帧时序上下文与视图条件路由,稳定预训练模型生成的静态结构。
  • 在AI City Challenge Track 5上达到顶尖团队水平,显著减少几何漂移。
  • 无需重训练,直接部署于预训练生成模型,适合实际交通系统应用。

长时程未来帧预测对自动驾驶、交通监控和智能交通系统至关重要,但受限于时间伪影、几何漂移和物体运动不一致等问题。现有潜空间视频扩散模型虽视觉质量出色,但在结构化交通场景中常导致几何不稳定与时序连贯性下降。本文提出一种无需训练的推理框架,通过多帧时序上下文与视图条件路由,稳定预训练视频生成结果中的可靠静态结构。针对前视摄像头视频,采用多帧深度分层渲染器,将观测历史帧中的静态几何投影至生成未来帧,同时保留生成模型中的动态区域。对于异构交通视角,使用冻结的视觉-语言模型从观测片段推断粗粒度摄像机组,并选择专用的基于运动的预测器。该框架无需重新训练或微调底层视频模型,可直接应用于预训练生成器。我们在AI City Challenge Track 5基准上验证了该方法,最终系统表现媲美顶尖队伍。结果表明,几何感知的推理时优化与视图条件的混合推理可在不修改原始模型架构的前提下,有效提升静态几何稳定性与低层结构保真度。

原文摘要 · Abstract (English)

Long-horizon future-frame prediction is important for autonomous driving, traffic surveillance, and intelligent transportation systems, yet remains challenging due to temporal ghosting, geometry drift, and inconsistent object motion. Recent latent video diffusion models have achieved impressive visual quality, but directly applying them to structured traffic scenes often leads to unstable geometry and degraded temporal coherence over extended horizons. We present a training-free inference framework that stabilizes reliable static structure in pretrained video predictions through multi-frame temporal context and view-conditioned routing. For front-camera videos, our method refines generated futures with a multi-frame depth-layered renderer that projects static geometry from observed history frames while preserving dynamic regions from the generative base model. For heterogeneous traffic views, a frozen vision-language model infers a coarse camera group from the observed clip and selects a specialized motion-based predictor. The framework requires neither retraining nor fine-tuning of the underlying video model and can be applied directly to pretrained generators. We validate the proposed framework on the AI City Challenge Track 5 benchmark, where our final system achieves competitive performance among the top-ranked teams. These results demonstrate that geometry-aware inference-time refinement and view-conditioned hybrid inference can improve static-geometry stability and low-level structural fidelity without changing the original model architecture.

交通预测视频生成几何建模推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。