arXiv:2608.05728cs.CVcs.LG2026-08

用参考帧生成视频,通过激活视觉记忆实现高精度重建。

Engram-E2VID: Reference-Based Event-to-Video Reconstruction via Generative Activation of Appearance Engrams

论文配图:Engram-E2VID: Reference-Based Event-to-Video Reconstruction via Generative Activation of Appearance Engrams
图 1 · 摘自论文原文
  • 将参考帧编码为视觉记忆,事件流构建运动结构骨架
  • 在扩散模型中逐步激活相关视觉记忆,提升重建质量
  • 适合长时序、复杂运动场景下的视频重建任务

基于参考帧的事件到视频重建旨在从参考帧和参考至目标时段内的事件流中恢复目标RGB帧。尽管事件提供精细时间线索,但其仅编码稀疏且异步的对数强度变化而非绝对外观,使忠实重建极具挑战性。核心难点在于将事件推导的目标时刻结构与参考帧中的相关外观信息关联起来,尤其在复杂运动和长时序间隔下。本文提出Engram-E2VID,一种结构引导框架,通过生成激活外观记忆实现目标帧重建。具体而言,参考帧被编码为令牌空间的外观记忆,事件流与参考上下文则转化为捕捉运动边界和事件引发结构变化的目标时刻运动-结构骨架。在单步扩散主干中,骨架衍生的结构令牌逐层与相关外观记忆交互并激活。该令牌空间关联使目标结构无需依赖像素级对应即可获取参考外观,而扩散先验则补充不确定或新出现区域。在三个基准测试中,Engram-E2VID相比最强同输入基线,峰值信噪比(PSNR)最高提升3.29 dB,学习感知图像相似性(LPIPS)降低0.08,且随重建间隔增长退化更缓慢。

原文摘要 · Abstract (English)

Reference-based event-to-video reconstruction aims to recover target RGB frames from a reference frame and the event stream captured over the reference-to-target interval. Although events provide fine-grained temporal cues, they encode sparse and asynchronous log-intensity changes rather than absolute appearance, making faithful reconstruction intrinsically challenging. The central challenge lies in associating event-derived target-time structures with relevant appearance information from the reference frame, especially under complex motion and long temporal intervals. In this work, we propose Engram-E2VID, a structure-guided framework that reconstructs target frames through the generative activation of appearance engrams. Specifically, the reference frame is encoded into token-space appearance engrams, while the event stream and reference context are transformed into a target-time motion-structure scaffold that captures motion boundaries and event-induced structural changes. Within a one-step diffusion backbone, scaffold-derived structural tokens progressively interact with and activate relevant appearance engrams across layers. This token-space association allows target structures to access reference appearance without relying on direct pixel-wise correspondence, while the diffusion prior complements uncertain or newly revealed regions. Across three benchmarks, Engram-E2VID improves PSNR by up to 3.29 dB and reduces LPIPS by up to 0.08 over the strongest same-input baseline, while degrading more slowly as the reconstruction interval increases.

事件相机视频生成扩散模型记忆机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。