首个基于Transformer的事件相机显著性预测模型,突破数据稀缺瓶颈。
Exploring deep learning for Event-Based Saliency Prediction with a Transformer-based model

- 用自监督预训练+轻量CNN解码器生成动态显著图
- 在合成数据上训练后仍可零样本迁移到真实事件流
- 构建两个新数据集,推动事件视觉注意力研究
显著性预测在RGB图像和视频中已有广泛研究,作为人类视觉注意的计算模型。相比之下,从事件数据中预测显著性仍基本未被探索,尽管事件相机具有生物启发性和优越的感知特性。两大障碍限制了该方向:缺乏大规模事件显著性数据集,以及缺少强基线方法。本文提出SEST(Swin Event-based Saliency Transformer),一种基于Transformer的事件数据显著性预测模型,通过事件原生预训练与合成监督克服数据稀缺问题。SEST采用自监督预训练的事件版Swin Transformer骨干网络,搭配轻量CNN解码器生成动态显著图。为缓解标注事件数据不足,我们构建了两个新基准数据集N-DHF1K和N-UCF Sports,源自大规模RGB显著性基准。实验表明,SEST显著优于现有事件显著性方法,并缩小了与先进RGB模型的性能差距。在真实事件相机数据集上的零样本评估进一步证明,该模型在合成数据上训练后仍具备跨域迁移能力。据我们所知,这是首次将深度学习应用于事件显著性预测的工作,开启了事件视觉与类脑视觉注意交叉研究的新方向。
原文摘要 · Abstract (English)
Saliency prediction has been extensively studied in RGB images and videos as a computational model of human visual attention. In contrast, predicting saliency from event-based data remains largely unexplored, despite the biological inspiration and favorable sensing properties of event cameras. Two obstacles have held this direction back: the absence of large-scale event saliency datasets, and the lack of a strong baseline. In this paper, we introduce SEST (Swin Event-based Saliency Transformer), a transformer-based model for saliency prediction from event data, bridging the data scarcity barrier through event-native pretraining and synthetic supervision. SEST leverages a self-supervised pretrained event-based Swin Transformer backbone combined with a lightweight CNN decoder to produce dynamic saliency maps. To address the scarcity of annotated event-based saliency data, we introduce two new benchmark datasets, N-DHF1K and N-UCF Sports, generated from large-scale RGB saliency benchmarks. Experimental results show that SEST clearly outperforms existing event-based saliency methods and narrows the performance gap with state-of-the-art RGB models. Zero-shot evaluation on a real event camera dataset further demonstrates that our model trained on synthetic data remains transferable on real event streams. To the best of our knowledge, this work is the first to apply deep learning to event-based saliency prediction, opening a new research direction at the intersection of event-based vision and neuromorphic visual attention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。