用单步扩散模型从事件流重建高质量彩色视频
EvDiff: Event-Based Video Reconstruction using One-Step Diffusion Models
- 设计单步扩散模型+时序一致编码器,降低高帧率生成计算开销
- 在真实数据集上实现像素级与感知质量双提升,优于现有方法
- 无需成对事件-图像数据,可利用大规模图像数据训练
类脑传感器事件相机以异步方式记录亮度变化,生成具有高时间分辨率和高动态范围的稀疏事件流。从事件流重建强度图像是一项高度病态的任务,因绝对亮度存在固有模糊性。早期方法通常采用端到端回归范式,直接确定性地将事件映射为强度帧。尽管在一定程度上有效,但这些方法常产生感知质量较差的结果,且难以扩展模型容量和训练数据规模。本文提出 EvDiff,一种基于代理训练框架的事件基扩散模型,可生成高质量视频。为降低高帧率视频生成的高计算成本,我们设计了仅执行一次前向扩散步骤的事件基扩散模型,并配备时序一致的 EvEncoder。此外,提出的新型代理训练框架消除了对成对事件-图像数据集的依赖,使模型能利用大规模图像数据集提升容量。所提 EvDiff 能够仅从单色事件流生成高质量彩色视频。在真实世界数据集上的实验表明,该方法在保真度与真实感之间取得良好平衡,在像素级和感知指标上均优于现有方法。代码将公开发布。
原文摘要 · Abstract (English)
As neuromorphic sensors, event cameras asynchronously record changes in brightness as streams of sparse events with the advantages of high temporal resolution and high dynamic range. Reconstructing intensity images from events is a highly ill-posed task due to the inherent ambiguity of absolute brightness. Early methods generally follow an end-to-end regression paradigm, directly mapping events to intensity frames in a deterministic manner. While effective to some extent, these approaches often yield perceptually inferior results and struggle to scale up in model capacity and training data. In this work, we propose EvDiff, an event-based diffusion model that follows a surrogate training framework to produce high-quality videos. To reduce the high computational cost of high-frame-rate video generation, we design an event-based diffusion model that performs only a single forward diffusion step, equipped with a temporally consistent EvEncoder. Furthermore, our novel Surrogate Training Framework eliminates the dependence on paired event-image datasets, allowing the model to leverage large-scale image datasets for higher capacity. The proposed EvDiff is capable of generating high-quality colorful videos solely from monochromatic event streams. Experiments on real-world datasets demonstrate that our method strikes a sweet spot between fidelity and realism, outperforming existing approaches on both pixel-level and perceptual metrics. The code will be released publicly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。