arXiv:2603.23487cs.CV2026-03被引 1

用25分钟真实数据训练事件相机运动估计,效果超群且适配视频插帧。

TETO: Tracking Events with Teacher Observation for Motion Estimation and Frame Interpolation

  • 通过教师-学生框架,从少量真实事件数据中学习运动估计。
  • 在EVIMO2和DSEC上达到顶尖追踪与光流性能,仅需极少训练数据。
  • 将精确运动信息用于视频插帧,显著提升生成质量,适合视觉算法研究者。

事件相机以微秒级精度捕捉像素亮度变化,提供RGB帧间丢失的连续运动信息。然而,现有基于事件的运动估计算法依赖大规模合成数据,常存在显著的仿真到现实差距。我们提出TETO(基于教师观察的事件追踪),一种通过知识蒸馏从预训练RGB追踪器学习的师生框架,仅需约25分钟未标注的真实世界录制数据。通过运动感知的数据筛选与查询采样策略,有效分离物体运动与主导的自我运动,最大化有限数据的学习效率。所获运动估计器可联合预测点轨迹与密集光流,作为显式运动先验,驱动预训练视频扩散变换器实现帧插值。在EVIMO2上实现最先进的点追踪,在DSEC上获得优异光流表现,训练数据量仅为以往的数个数量级;并验证了精准运动估计直接带来BS-ERGB与HQ-EVFI上更优的帧插值效果。

原文摘要 · Abstract (English)

Event cameras capture per-pixel brightness changes with microsecond resolution, offering continuous motion information lost between RGB frames. However, existing event-based motion estimators depend on large-scale synthetic data that often suffers from a significant sim-to-real gap. We propose TETO (Tracking Events with Teacher Observation), a teacher-student framework that learns event motion estimation from only $\sim$25 minutes of unannotated real-world recordings through knowledge distillation from a pretrained RGB tracker. Our motion-aware data curation and query sampling strategy maximizes learning from limited data by disentangling object motion from dominant ego-motion. The resulting estimator jointly predicts point trajectories and dense optical flow, which we leverage as explicit motion priors to condition a pretrained video diffusion transformer for frame interpolation. We achieve state-of-the-art point tracking on EVIMO2 and optical flow on DSEC using orders of magnitude less training data, and demonstrate that accurate motion estimation translates directly to superior frame interpolation quality on BS-ERGB and HQ-EVFI.

事件相机运动估计视频插帧知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。