arXiv:2605.07910cs.CV2026-05

解决车路协同中时间不同步导致的动态物体重影问题。

One World, Dual Timeline: Decoupled Spatio-Temporal Gaussian Scene Graph for 4D Cooperative Driving Reconstruction

论文配图:One World, Dual Timeline: Decoupled Spatio-Temporal Gaussian Scene Graph for 4D Cooperative Driving Reconstruction
图 1 · 摘自论文原文
  • 为车辆与路侧设备分别建立独立时间轨迹,共享外观高斯点集。
  • 在26段V2X-Seq数据上,动态区域PSNR提升3.2dB,视频质量提升37.7%。
  • 适合车路协同4D重建、对时间异步敏感的自动驾驶场景。

从车路协同自动驾驶(VICAD)数据重建动态场景面临核心挑战:车辆与路侧摄像头使用独立时钟,同一动态目标如车辆和行人被不同源以不同物理时间捕获。现有高斯场景图方法隐含同步观测假设,为每帧每个目标分配单一姿态,该假设在协同设置下失效,引发梯度冲突并导致动态目标严重重影。我们证明这是表征层面的根本缺陷,而非优化误差:任何单时间线模型在目标速度与跨源时间偏移下均存在不可消除的光度损失,其量级随速度和偏移平方增长。为此,提出DUST(Decoupled Spatio-Temporal)高斯场景图用于4D车路协同重建。DUST为每个目标共享统一的高斯点集以保证外观一致性,同时维护与各源真实拍摄时间戳对齐的独立姿态轨迹。我们证明该解耦使姿态梯度核呈块对角化,彻底消除跨源干扰。为提升实用性,进一步引入基于静态锚点的姿态校正流程,修复车辆与路侧标注间的空间错位,并设计姿态正则化的联合优化方案,防止训练初期轨迹抖动与漂移。在V2X-Seq的26个序列上,DUST实现当前最优性能,动态区域PSNR较最强基线提升3.2 dB,Fréchet Video Distance降低37.7%,且在更大时间异步条件下仍保持鲁棒性。

原文摘要 · Abstract (English)

Reconstructing dynamic scenes from Vehicle-to-Infrastructure Cooperative Autonomous Driving (VICAD) data is fundamentally complicated by temporal asynchrony: vehicle and infrastructure cameras operate on independent clocks, capturing the same dynamic agent such as cars and pedestrians at different physical times. Existing Gaussian Scene Graph methods implicitly assume synchronized observations and assign a single pose per agent per frame, which is an assumption that breaks in cooperative settings, where the resulting gradient conflicts cause severe ghosting on dynamic agents. We identify this as a representation-level failure, not an optimization artifact: we prove that any single-timeline formulation incurs an irreducible photometric loss scaling quadratically with agent velocity and cross-source time offset. To resolve this, we propose Dust (DecoUpled Spatio-Temporal) Gaussian Scene Graph for 4D Cooperative Driving Reconstruction. DUST Gaussian Scene Graph shares a canonical Gaussian set per agent for appearance consistency, while maintaining decouple pose trajectories aligned to each source's true capture timestamps. We prove that this decoupling enables the pose-gradient kernel block-diagonal, eliminating cross-source interference entirely. To make Dust practical, we further introduce a static anchor-based pose correction pipeline that corrects spatio misalignment between vehicle and infrastructure annotations, and a pose-regularized joint optimization scheme that prevents trajectory jitter and drift during early training. On 26 sequences from V2X-Seq, DUST achieves state-of-the-art performance, improving dynamic-area PSNR by 3.2 dB over the strongest baseline and reducing Fréchet Video Distance by 37.7%, with keeping robustness under larger temporal asynchrony.

4D重建车路协同高斯建模时间异步

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。