arXiv:2510.17347cs.CV2025-10

用语义信息提升事件相机视频重建质量,让画面更真实。

Semantic-E2VID: a Semantic-Enriched Paradigm for Event-to-Video Reconstruction

  • 将事件流与分割模型语义融合,增强物体结构理解
  • 在6个公开数据集上超越现有方法,重建更清晰
  • 适合做事件视觉重建、自动驾驶感知的研究者

事件相机通过异步记录亮度变化,为高速高动态范围视觉提供潜力。事件到视频(E2V)重建是事件视觉的基础任务,旨在从事件流恢复强度视频。现有方法多将重建视为时空信号恢复问题,依赖时间聚合和空间特征学习推断帧。然而,事件数据因变化驱动机制,本质上缺乏对象级结构和上下文信息,导致重建不准确。本文提出语义增强型端到端E2V框架Semantic-E2VID,从语义视角重构问题,强调显式补全缺失语义信息的重要性。首先,通过预训练分割模型SAM,将事件表示与语义对齐,避免模态漂移;其次,以表示兼容方式将语义融合进事件隐空间,使事件特征具备对象结构与上下文线索;最后引入语义感知监督,指导重建聚焦于语义有意义区域,补充像素级与时间目标。六个公开基准测试表明,Semantic-E2VID持续优于当前最优方法。

原文摘要 · Abstract (English)

Event cameras provide a promising sensing modality for high-speed and high-dynamic-range vision by asynchronously capturing brightness changes. A fundamental task in event-based vision is event-to-video (E2V) reconstruction, which aims to recover intensity videos from event streams. Most existing E2V approaches formulate reconstruction as a temporal--spatial signal recovery problem, relying on temporal aggregation and spatial feature learning to infer intensity frames. While effective to some extent, this formulation overlooks a critical limitation of event data: due to the change-driven sensing mechanism, event streams are inherently semantically under-determined, lacking object-level structure and contextual information that are essential for faithful reconstruction. In this work, we revisit E2V from a semantic perspective and argue that effective reconstruction requires going beyond temporal and spatial modeling to explicitly account for missing semantic information. Based on this insight, we propose \textit{Semantic-E2VID}, a semantic-enriched end-to-end E2V framework that reformulates reconstruction as a process of semantic learning, fusing and decoding. Our approach first performs semantic abstraction by bridging event representations with semantics extracted from a pretrained Segment Anything Model (SAM), while avoiding modality-induced feature drift. The learned semantics are then fused into the event latent space in a representation-compatible manner, enabling event features to capture object-level structure and contextual cues. Furthermore, semantic-aware supervision is introduced to explicitly guide the reconstruction process toward semantically meaningful regions, complementing conventional pixel-level and temporal objectives. Extensive experiments on six public benchmarks demonstrate that Semantic-E2VID consistently outperforms state-of-the-art E2V methods.

事件视觉语义增强视频重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。