arXiv:2509.22864cs.CV2025-09中稿 · WACV2026被引 1

用图像生成模型生成可控事件数据,降低标注成本。

ControlEvents: Controllable Synthesis of Event Camera Datawith Foundational Prior from Image Diffusion Models

  • 利用图像扩散模型先验生成事件数据,支持文本、骨骼等多类控制信号。
  • 合成数据可提升视觉识别、姿态估计等任务性能,效果优于基线。
  • 支持未见文本标签生成,适合缺乏标注数据的事件视觉研究。

近年来,事件相机因其高时间分辨率和高动态范围等生物启发特性受到广泛关注。然而,获取大规模带标注的真实事件数据仍面临挑战且成本高昂。本文提出ControlEvents,一种基于扩散模型的生成方法,可生成高质量事件数据,并受类别文本标签、2D骨骼、3D身体姿态等多种控制信号引导。核心思路是利用如Stable Diffusion等基础模型的扩散先验,在极少微调和有限标注数据条件下实现高质量生成。该方法简化了数据生成流程,显著降低标注事件数据集的成本。实验表明,合成的标注事件数据能有效提升视觉识别、2D骨骼估计和3D身体姿态估计任务的模型性能。此外,该方法在训练时未见过的文本标签下也能生成事件数据,展现出从基础模型继承的强大文本生成能力。

原文摘要 · Abstract (English)

In recent years, event cameras have gained significant attention due to their bio-inspired properties, such as high temporal resolution and high dynamic range. However, obtaining large-scale labeled ground-truth data for event-based vision tasks remains challenging and costly. In this paper, we present ControlEvents, a diffusion-based generative model designed to synthesize high-quality event data guided by diverse control signals such as class text labels, 2D skeletons, and 3D body poses. Our key insight is to leverage the diffusion prior from foundation models, such as Stable Diffusion, enabling high-quality event data generation with minimal fine-tuning and limited labeled data. Our method streamlines the data generation process and significantly reduces the cost of producing labeled event datasets. We demonstrate the effectiveness of our approach by synthesizing event data for visual recognition, 2D skeleton estimation, and 3D body pose estimation. Our experiments show that the synthesized labeled event data enhances model performance in all tasks. Additionally, our approach can generate events based on unseen text labels during training, illustrating the powerful text-based generation capabilities inherited from foundation models.

事件相机扩散模型数据生成可控合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。