用事件相机数据生成高精度未来运动轨迹,提升视觉系统预测能力。
E-Motion: Future Motion Simulation via Event Sequence Diffusion
- 结合扩散模型与事件序列,实现细粒度运动模拟
- 通过强化学习对齐逆向生成轨迹,提升预测准确性
- 适用于自动驾驶与机器人导航等实时交互场景
预测物体未来运动是计算机视觉中理解与交互动态环境的关键任务。基于事件的传感器能以极高的时间分辨率捕捉场景变化,为实现前所未有的细节与精度的运动预测提供了可能。受此启发,我们提出将视频扩散模型的强大学习能力与事件相机丰富的运动信息相结合,构建运动模拟框架。具体而言,首先利用预训练的稳定视频扩散模型适配事件序列数据集,促进从RGB视频到事件主导领域的知识迁移。此外,引入基于强化学习的对齐机制,优化扩散模型的逆向生成轨迹,显著提升性能与准确性。在多种复杂场景下的大量测试验证了该方法的有效性,展现出在自动驾驶引导、机器人导航和交互媒体等应用中革新运动流预测的潜力。研究结果为提升计算机视觉系统的解释力与预测精度指明了新方向。
原文摘要 · Abstract (English)
Forecasting a typical object's future motion is a critical task for interpreting and interacting with dynamic environments in computer vision. Event-based sensors, which could capture changes in the scene with exceptional temporal granularity, may potentially offer a unique opportunity to predict future motion with a level of detail and precision previously unachievable. Inspired by that, we propose to integrate the strong learning capacity of the video diffusion model with the rich motion information of an event camera as a motion simulation framework. Specifically, we initially employ pre-trained stable video diffusion models to adapt the event sequence dataset. This process facilitates the transfer of extensive knowledge from RGB videos to an event-centric domain. Moreover, we introduce an alignment mechanism that utilizes reinforcement learning techniques to enhance the reverse generation trajectory of the diffusion model, ensuring improved performance and accuracy. Through extensive testing and validation, we demonstrate the effectiveness of our method in various complex scenarios, showcasing its potential to revolutionize motion flow prediction in computer vision applications such as autonomous vehicle guidance, robotic navigation, and interactive media. Our findings suggest a promising direction for future research in enhancing the interpretative power and predictive accuracy of computer vision systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。