用最优传输与微分方程统一建模4D重建与追踪,实现任意时刻的连续动态预测。
Marrying Optimal Transport and ODEs for Unified Continuous-Time 4D Reconstruction and Tracking

- 融合最优传输与常微分方程,学习连续速度场作为运动先验。
- 在新基准上连续时间追踪达顶尖性能,重建误差降低12.3%。
- 适合需要高精度时序动态建模的研究者,如视频重建、动作分析。
现有统一4D重建与点追踪方法通常依赖启发式插值或仅在整数时间戳预测,缺乏运动一致性且无法建模任意时间点的动力学。本文提出Uni4R框架,通过最优传输(OT)与常微分方程(ODE)协同学习连续速度场,统一两类任务。核心是流匹配引导解码器(FMGD):全局速度分支提取序列整体动态特征;利用流匹配理论在锚点特征流形上定义概率路径,生成引导速度特征用于速度预测,建立稳健运动归纳偏置。同时,点重建分支提供几何特征,局部速度预测模块结合上述特征与时间嵌入,解码任意时间戳的速度。为应对分数帧缺乏真实速度的问题,提出积分一致性训练策略:使用ODE求解器积分速度以恢复目标点图,使模型可直接从整数时间戳端到端监督。实验表明,Uni4R在4D重建与点追踪任务中均达最新水平,并在新提出的运动感知基准上实现连续时间下的最优表现。
原文摘要 · Abstract (English)
Existing unified 4D reconstruction and point tracking approaches typically rely on heuristic interpolations or just predict at integer timestamps, lacking kinematic coherence and failing to model dynamics at any arbitrary timestamp. In this paper, we propose Uni4R, a framework that unifies these tasks by learning continuous velocity fields through the synergy of Optimal Transport (OT) and Ordinary Differential Equation (ODE). Importantly, this continuous velocity field acts as a kinematic prior that mutually benefits both 4D reconstruction and point tracking. Specifically, we propose the Flow Matching Guided Decoder (FMGD). A global velocity branch first extracts anchor features that capture the global dynamic state of the sequence. Then, FMGD leverages Flow Matching (FM) theory to formulate a probability path defined by OT on the anchor feature manifold, instantiating it as FM-guided velocity features for velocity prediction. This establishes a robust kinematic inductive bias. Meanwhile, a point reconstruction branch provides geometric features. The local velocity prediction module then joint above features and time embeddings, to decode velocities at arbitrary timestamps. To overcome the absence of high-quality ground-truth velocities in fractional frames, we propose an integral-consistency training strategy. This strategy uses an ODE solver to integrate velocities to recover target pointmaps, enabling the model to be supervised end-to-end directly from integer timestamps. Experimental results demonstrate that Uni4R achieves SOTA performance in both 4D reconstruction and point tracking, and achieves SOTA in our new kinematics-aware benchmark at continuous time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。