arXiv:2510.12753cs.CV2025-10NeurIPS被引 4

用隐式正则化联合估计事件相机的运动与光流,无需监督。

E-MoFlow: Learning Egomotion and Optical Flow from Event Data via Implicit Regularization

  • 通过隐式神经表示和样条建模,自然融合时空一致性。
  • 在无监督下实现最优性能,超越现有自监督方法。
  • 适合事件相机、自动驾驶等需要鲁棒运动估计场景。

光学流与6-自由度自运动估计是3D视觉中的基础任务,传统上被独立处理。对于类脑视觉(如事件相机),由于缺乏可靠的特征匹配,在无真值监督下独立求解二者会带来病态问题。现有方法通过显式变分正则化或结构-运动先验参数化来缓解,前者引入偏差并增加计算开销,后者常陷入次优局部极小。为此,本文提出一种无监督框架E-MoFlow,通过隐式时空与几何正则化联合优化自运动与光流。首先,将相机运动建模为连续样条,光流采用隐式神经表示,利用归纳偏置自然嵌入时空一致性;其次,通过微分几何约束融入结构-运动先验,避免显式深度估计的同时保持严格的几何一致性。实验表明,该框架适用于一般6-自由度运动场景,在无监督方法中达到最先进水平,甚至媲美有监督方法。

原文摘要 · Abstract (English)

The estimation of optical flow and 6-DoF ego-motion, two fundamental tasks in 3D vision, has typically been addressed independently. For neuromorphic vision (e.g., event cameras), however, the lack of robust data association makes solving the two problems separately an ill-posed challenge, especially in the absence of supervision via ground truth. Existing works mitigate this ill-posedness by either enforcing the smoothness of the flow field via an explicit variational regularizer or leveraging explicit structure-and-motion priors in the parametrization to improve event alignment. The former notably introduces bias in results and computational overhead, while the latter, which parametrizes the optical flow in terms of the scene depth and the camera motion, often converges to suboptimal local minima. To address these issues, we propose an unsupervised framework that jointly optimizes egomotion and optical flow via implicit spatial-temporal and geometric regularization. First, by modeling camera's egomotion as a continuous spline and optical flow as an implicit neural representation, our method inherently embeds spatial-temporal coherence through inductive biases. Second, we incorporate structure-and-motion priors through differential geometric constraints, bypassing explicit depth estimation while maintaining rigorous geometric consistency. As a result, our framework (called E-MoFlow) unifies egomotion and optical flow estimation via implicit regularization under a fully unsupervised paradigm. Experiments demonstrate its versatility to general 6-DoF motion scenarios, achieving state-of-the-art performance among unsupervised methods and competitive even with supervised approaches.

事件相机自运动估计光流无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。