揭示事件相机多模态运动估计的隐空间几何结构
On the Geometry of Learned Representations in Event-Based Multi-Modal Egomotion Estimation

- 用跨模态注意力融合事件、惯性与深度信号
- 嵌入向量位于与运动变量对齐的低维流形上
- 结果连接传统理论与现代神经网络方法
传统事件相机自运动估计方法依赖对比度最大化、单应性估计或密集光流结合解析运动反解等几何优化框架。本文研究多模态网络中学习到的表示几何结构。通过跨模态注意力架构融合事件张量、惯性测量与距离信号,并在批次设置下训练。分析发现:(i)嵌入位于与运动变量对齐的低维流形上;(ii)注意力权重随角度激励和视觉可靠性动态调整;(iii)融合表示恢复了经典可观测性线索。这些结果将解析估计理论与数据驱动融合方法相连接。
原文摘要 · Abstract (English)
Classical approaches to event-based egomotion estimation, including those adopted by the top-performing teams of the ELOPE challenge, rely on geometric optimization frameworks such as contrast maximization, homography estimation, or dense optical flow combined with analytic motion inversion. This work investigates the geometric structure that emerges inside a multi-modal network for egomotion estimation. Event tensors, inertial measurements, and range signals are fused through a cross-modal attention architecture and trained in a batch setting. We analyze the latent space geometry and attention dynamics, showing that (i) embeddings lie on low-dimensional manifolds aligned with motion variables, (ii) attention weights adapt with angular excitation and visual reliability, and (iii) the fused representation recovers classical observability cues. These results bridge analytical estimation theory and modern data-driven fusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。