arXiv:2605.19329cs.CVcs.AI2026-05

用事件相机提升视觉语言模型在恶劣环境下的理解能力

RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding

论文配图:RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding
图 1 · 摘自论文原文
  • 双流架构融合RGB图像与事件流,同步处理视觉信息
  • 在光照不足等挑战场景下,图文生成与问答性能显著优于单一模态模型
  • 构建了两个新数据集,支持复杂环境下的视觉语言任务研究

传统视觉语言模型在低光、高动态范围或快速运动等恶劣环境下表现不佳,因标准RGB图像在此类条件下质量下降。事件相机以高时间分辨率和宽动态范围异步记录像素亮度变化,能有效保留运动信息。本文提出RE-VLM,首个联合使用RGB图像与事件流的双流视觉语言模型,实现对正常及挑战性场景的鲁棒理解。模型采用并行的RGB与事件编码器,并设计渐进式训练策略对齐异构视觉特征与语言表示。为应对RGB-事件-文本标注数据稀缺问题,提出基于图的流水线,将同步的RGB-事件流转化为可验证场景图,进而合成描述与问答对。为评估模型,构建两个数据集:PEOD-Chat(聚焦光照挑战场景)与RGBE-Chat(覆盖多样化场景)。在图文生成与视觉问答基准上,RE-VLM在参数量相当的情况下,持续优于当前主流的纯RGB或纯事件模型,尤其在困难条件下提升显著。结果证明事件增强型视觉语言模型可在多种真实环境中实现更鲁棒的视觉语言理解。

原文摘要 · Abstract (English)

Conventional vision-language models (VLMs) struggle to interpret scenes captured under adverse conditions (e.g., low light, high dynamic range, or fast motion) because standard RGB images degrade in such environments. Event cameras provide a complementary modality: they asynchronously record per-pixel brightness changes with high temporal resolution and wide dynamic range, preserving motion cues where frames fail. We propose RE-VLM, the first dual-stream vision-language model that jointly leverages RGB images and event streams for robust scene understanding across both normal and challenging conditions. RE-VLM employs parallel RGB and event encoders together with a progressive training strategy that aligns heterogeneous visual features with language. To address the scarcity of RGB-Event-Text supervision, we further propose a graph-driven pipeline that converts synchronized RGB-Event streams into verifiable scene graphs, from which we synthesize captions and question-answer (QA) pairs. To develop and evaluate RE-VLM, we construct two datasets: PEOD-Chat, targeting illumination-challenged scenes, and RGBE-Chat, covering diverse scenarios. On captioning and VQA benchmarks, RE-VLM consistently outperforms state-of-the-art RGB-only and event-only models with comparable parameter counts, with particularly large gains under challenging conditions. These results demonstrate the effectiveness of event-augmented VLMs in achieving robust vision-language understanding across a wide range of real-world environments.

视觉语言模型事件相机多模态学习鲁棒理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。