MotionSync让因果追踪与非因果修正共用一套系统,提升3D感知效率。
MotionSync: Non-Causal Refinement of Causal Tracker for Label-Efficient 3D Perception

- 在因果追踪基础上,引入不确定性校准和多假设运动模型
- 非因果阶段用平滑与语义剪枝提升轨迹质量,增益达+3.3 mAP/L2
- 适合需兼顾实时与离线场景的自动驾驶数据标注任务
三维框与轨迹标注是自动驾驶数据引擎的成本瓶颈。现有离线系统完全替代在线感知,导致团队需维护两套系统。MotionSync将因果/非因果边界显式设计为架构分界。一个严格因果的追踪器基于强基线,结合创新性不确定性校准、帧率无关运动关联门控及学习模式选择的多假设运动模型,输出有效在线结果。随后的非因果阶段对缓存轨迹进行独立的姿态、尺度、航向平滑,物理验证的间隙补全,以及基于激光雷达点云标签的鬼影轨迹剔除。修正器不回写,因此同一系统同时服务两种模式,且优化效果仅作为因果估计的增量。作为自动标注器,仅用25%人工标签加MotionSync伪标签训练的固定3D检测器,在Waymo上达到全监督模型96.9%的mAP;在10%预算下,非因果阶段较同追踪器因果阶段伪标签提升+3.3 mAP/L2。将在线追踪器重新拟合其自身修正输出,可恢复73%人类标注的收益,而原始因果输出反而劣于不重拟合。作为追踪器,MotionSync在主流指标上与最优离线方法持平,并在误差组成上更优——正是优化可发挥作用之处:同时减少漏检与碎片化,体现间隙补全效果而非检测器调优。
原文摘要 · Abstract (English)
Three-dimensional box-and-track annotation is the cost bottleneck in autonomous-driving data engines, and the offline systems built to relieve it replace the online perception stack outright, so a team needing both regimes maintains and reconciles two. MotionSync makes the causal/non-causal boundary an explicit architectural seam instead. A strictly causal tracker, built on a strong published baseline and extended with innovation-driven uncertainty calibration, frame-rate-invariant kinematic association gates, and multi-hypothesis motion with learned mode selection, emits a valid online result. A non-causal pass then revises the buffered trajectories with Rauch--Tung--Striebel smoothing applied separately to pose, extent and yaw, physics-validated gap completion, and semantic pruning of ghost tracks against LiDAR point labels. The refiner never writes back, so one system serves both regimes and refinement's effect is a delta over an unaltered causal estimate. Used as an auto-labeller, a fixed 3D detector trained on 25% human labels plus MotionSync pseudo-labels reaches 96.9% of its full-supervision mean average precision (mAP) on Waymo, and at a 10% budget the non-causal pass accounts for +3.3 mAP/L2 over pseudo-labels from the same tracker's causal stage. Re-fitting the online tracker on its own refined output recovers 73% of the benefit of human supervision, while its causal output is worse supervision than no re-fitting at all. As a tracker MotionSync is at parity with the leading published offline entries on the headline metric and ahead of them on error composition, which is where a refinement pass can act at all: it reduces misses and fragmentations together, the signature of gap completion rather than of a tuned detector.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。