用视频扩散模型从稀疏事件数据重建高清视频帧
UniE2F: A Unified Diffusion Framework for Event-to-Frame Reconstruction with Video Foundation Models
- 利用预训练视频扩散模型的生成先验,将事件数据转为视频帧
- 事件帧间残差引导提升重建精度,实现更真实纹理还原
- 零样本支持插帧与预测,统一处理事件到帧的任务
事件相机在高速、低功耗和高动态范围场景感知中表现优异,但因其仅记录相对亮度变化而非绝对强度,导致数据流丢失大量空间信息和静态纹理。本文通过利用预训练视频扩散模型的生成先验,从稀疏事件数据中重建高保真视频帧。首先构建基线模型,直接以事件数据为条件生成视频;随后基于事件流与视频帧间的物理关联,引入事件驱动的帧间残差引导,提升重建准确性。进一步地,通过调节反向扩散采样过程,零样本扩展至视频帧插值与预测任务,形成统一的事件到帧重建框架。在真实世界与合成数据集上的实验表明,本方法在定量与定性评价上均显著优于现有方法。视频演示见补充材料,代码将公开于https://github.com/CS-GangXu/UniE2F。
原文摘要 · Abstract (English)
Event cameras excel at high-speed, low-power, and high-dynamic-range scene perception. However, as they fundamentally record only relative intensity changes rather than absolute intensity, the resulting data streams suffer from a significant loss of spatial information and static texture details. In this paper, we address this limitation by leveraging the generative prior of a pre-trained video diffusion model to reconstruct high-fidelity video frames from sparse event data. Specifically, we first establish a baseline model by directly applying event data as a condition to synthesize videos. Then, based on the physical correlation between the event stream and video frames, we further introduce the event-based inter-frame residual guidance to enhance the accuracy of video frame reconstruction. Furthermore, we extend our method to video frame interpolation and prediction in a zero-shot manner by modulating the reverse diffusion sampling process, thereby creating a unified event-to-frame reconstruction framework. Experimental results on real-world and synthetic datasets demonstrate that our method significantly outperforms previous approaches both quantitatively and qualitatively. We also refer the reviewers to the video demo contained in the supplementary material for video results. The code will be publicly available at https://github.com/CS-GangXu/UniE2F.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。