arXiv:2503.18552cs.CVcs.AI2025-03被引 3

用事件相机数据生成更精准的人体动画,抗模糊低光效果好。

EvAnimate: Event-conditioned Image-to-Video Generation for Human Animation

  • 将异步事件流转为三通道表示,适配扩散模型生成
  • 在低光、运动模糊场景下动画质量显著提升
  • 适合做真实复杂环境下的动作生成研究

条件人体动画通常使用视频提取的姿势运动信息,但这类信息常存在时间分辨率低、运动模糊和光照变化下表现不可靠的问题。事件相机能提供高时间分辨率且对运动模糊、低光和曝光变化鲁棒的运动信息。本文提出 EvAnimate,首个利用事件流作为精确运动提示进行条件人体图像动画生成的方法。该方法通过自适应切片率与密度编码事件数据为专用三通道表示,与扩散生成模型完全兼容。采用双分支结构显式建模事件驱动动态,实现高质量且时间连贯的动画,在复杂现实条件下性能显著提升。通过特殊增强策略进一步提高跨主体泛化能力。为促进后续研究,我们构建新基准,包含模拟事件数据和真实世界事件数据集,覆盖正常与挑战性场景下的人体动作。实验表明,EvAnimate 在传统视频线索失效场景中仍保持高时间保真度与强鲁棒性。

原文摘要 · Abstract (English)

Conditional human animation traditionally animates static reference images using pose-based motion cues extracted from video data. However, these video-derived cues often suffer from low temporal resolution, motion blur, and unreliable performance under challenging lighting conditions. In contrast, event cameras inherently provide robust and high temporal-resolution motion information, offering resilience to motion blur, low-light environments, and exposure variations. In this paper, we propose EvAnimate, the first method leveraging event streams as robust and precise motion cues for conditional human image animation. Our approach is fully compatible with diffusion-based generative models, enabled by encoding asynchronous event data into a specialized three-channel representation with adaptive slicing rates and densities. High-quality and temporally coherent animations are achieved through a dual-branch architecture explicitly designed to exploit event-driven dynamics, significantly enhancing performance under challenging real-world conditions. Enhanced cross-subject generalization is further achieved using specialized augmentation strategies. To facilitate future research, we establish a new benchmarking, including simulated event data for training and validation, and a real-world event dataset capturing human actions under normal and challenging scenarios. The experiment results demonstrate that EvAnimate achieves high temporal fidelity and robust performance in scenarios where traditional video-derived cues fall short.

人体动画事件相机扩散模型生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。