arXiv:2603.03969cs.CV2026-03

用视觉模型蒸馏提升事件流表示能力,实现高分辨率下更精准的跨模态对齐。

Scaling Dense Event-Stream Pretraining from Visual Foundation Models

  • 通过结构感知蒸馏损失,强化图像与事件流间的语义对齐。
  • 在多个下游任务中显著优于传统方法,提升泛化性与数据效率。
  • 适合做事件流表征学习、智能视觉系统等研究方向的开发者。

从不规则事件流中学习通用且细粒度的表示至关重要,但受限于高昂标注成本,难以扩展数据规模、语义丰富度与应用范围。为此,我们提出一种新型自监督预训练方法,将视觉基础模型(VFMs)的知识蒸馏到事件流中,以规模化提升事件表示能力。具体而言,我们构建了一个大规模同步图像-事件数据集,增强跨模态对齐。然而,由于图像与事件在稀疏性和粒度上的固有差异,现有蒸馏范式容易导致事件表示语义坍塌,尤其在高分辨率下更为明显。为弥合此差距,我们扩展对齐目标至由VFMs提供的现成语义结构,提供更广感受野与更强监督。核心在于结构感知蒸馏损失,引导更高质量的图像-事件对应关系,优化密集事件表示。大量实验表明,该方法在下游基准上取得显著突破,显著超越传统方法与现有预训练技术,在泛化性、数据效率和可迁移性方面均有提升。

原文摘要 · Abstract (English)

Learning versatile, fine-grained representations from irregular event streams is pivotal yet nontrivial, primarily due to the heavy annotation that hinders scalability in dataset size, semantic richness, and application scope. To mitigate this dilemma, we launch a novel self-supervised pretraining method that distills visual foundation models (VFMs) to push the boundaries of event representation at scale. Specifically, we curate an extensive synchronized image-event collection to amplify cross-modal alignment. Nevertheless, due to inherent mismatches in sparsity and granularity between image-event domains, existing distillation paradigms are prone to semantic collapse in event representations, particularly at high resolutions. To bridge this gap, we propose to extend the alignment objective to semantic structures provided off-the-shelf by VFMs, indicating a broader receptive field and stronger supervision. The key ingredient of our method is a structure-aware distillation loss that grounds higher-quality image-event correspondences for alignment, optimizing dense event representations. Extensive experiments demonstrate that our approach takes a great leap in downstream benchmarks, significantly surpassing traditional methods and existing pretraining techniques. This breakthrough manifests in enhanced generalization, superior data efficiency and elevated transferability.

事件流自监督跨模态蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。