arXiv:2605.16420cs.CVcs.LG2026-05中稿 · the 1st Workshop o…

用轨迹引导扩散模型重建无人机拍摄的船舶航行视频缺失帧。

Video Reconstruction using Diffusion-based Image-to-Video Generation with Trajectory Guidance

论文配图:Video Reconstruction using Diffusion-based Image-to-Video Generation with Trajectory Guidance
图 1 · 摘自论文原文
  • 通过GPS投影生成运动线索,引导预训练图像转视频模型生成视频。
  • 生成帧在视觉自然度、运动平滑性和轨迹一致性上均优于基线方法。
  • 适用于低纹理、小目标的海上视频修复,无需领域微调。

本文针对自主水面航行器执行结构化海上操作时,俯视无人机视频中缺失或丢帧的问题提出解决方案。我们构建了一条流水线,仅需原始GPS遥测数据和单张参考帧,即可通过预训练的图像到视频扩散模型生成轨迹引导的视频序列,无需领域特定微调。船上遥测日志中的GPS坐标经等距圆柱投影映射至图像空间,生成每艘船的运动线索,作为SG-I2V扩散模型的条件输入。生成帧通过感知、时间与轨迹指标与真实视频对比,并与光流外推和RIFE插值基线方法进行比较。SG-I2V在所有方法中生成最自然的帧(BRISQUE 25.52,接近真实值23.64),运动幅度最真实(时间平滑度1.14,真实值1.42),且轨迹对齐最强(误差9.31像素,真实轨迹误差28.70像素,后者反映视频与GPS日志间近似的时间对齐而非生成误差),证明在低纹理、小目标条件下,轨迹引导的扩散合成是可行的海上视频重建方法。

原文摘要 · Abstract (English)

This paper addresses the problem of reconstructing missing or dropped frames in top-down drone video of autonomous surface vehicles performing structured maritime manoeuvres. We propose a pipeline that converts raw GPS telemetry and a single reference frame into a trajectory-guided video sequence using a pre-trained image-to-video diffusion model, requiring no domain-specific fine-tuning. GPS coordinates from onboard telemetry logs are projected into image space via an equirectangular mapping, producing per-vessel motion cues that condition the SG-I2V diffusion model. The generated frames are evaluated against ground-truth video using perceptual, temporal and trajectory-based metrics, and benchmarked against optical flow extrapolation and RIFE interpolation baselines. SG-I2V produces the most naturally appearing frames among all methods (BRISQUE 25.52, closest to ground-truth 23.64), the most realistic motion magnitude (temporal smoothness 1.14 vs. ground truth 1.42), and the strongest GPS trajectory adherence (9.31px vs. 28.70px for ground-truth, the latter reflecting approximate temporal alignment between footage and GPS logs rather than generation error), demonstrating that trajectory-guided diffusion synthesis is a viable approach to maritime video reconstruction under challenging low-texture, small-object conditions.

视频生成扩散模型轨迹引导视频修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。