arXiv:2607.08770cs.CV2026-07International Conf…

用扩散模型统一解决事件视频重建、预测与插帧,长期稳定且效果出色。

LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models

论文配图:LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models
图 1 · 摘自论文原文
  • 基于预训练视频扩散模型微调,联合处理三类任务
  • 在真实数据集上显著优于现有方法,长序列保持时序连贯
  • 适合需要高精度视频生成的机器人、自动驾驶场景

从稀疏事件流中恢复高质量视频是一项挑战性任务。回归方法常导致纹理模糊,而现有生成模型在长时间序列上难以保持稳定性。本文提出LongE2V,利用预训练视频扩散先验,联合实现事件驱动视频的重建、预测与帧插值。通过微调基础视频模型,本方法具备高数据效率和优异感知质量。引入自回归展开与自适应上下文切换以缓解极长序列中的时间漂移问题;提出重编码对齐与交叉残差校正,确保插帧过程双向一致性;采用事件体素密度增强提升不同传感器分辨率下的鲁棒性。在多个真实世界基准上的大量实验表明,LongE2V在三项任务上均超越当前最优方法,表现出卓越的时间连贯性与零样本泛化能力。

原文摘要 · Abstract (English)

Recovering high-quality video from sparse event streams is a challenging task. Regression methods often blur textures, while existing generative models struggle with long-term stability. We propose LongE2V, a novel approach that leverages pre-trained video diffusion priors to jointly handle event-based video reconstruction, prediction, and frame interpolation. By fine-tuning a foundational video model, our approach achieves high data efficiency and superior perceptual quality. We introduce Autoregressive Unrolling and Adaptive Context Switching to mitigate temporal drift in extremely long sequences. We also propose Reencoding Alignment with Cross Residual Correction to ensure precise bidirectional consistency during frame interpolation. Furthermore, Event Voxel Density Augmentation ensures robustness across varying sensor resolutions. Extensive experiments on real-world benchmarks demonstrate that LongE2V outperforms state-of-the-art methods across all three tasks, exhibiting exceptional temporal coherence and zero-shot generalization. Project page: https://cdfan0627.github.io/LongE2V-page/

事件视频扩散模型视频重建帧插值

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。