用视频追踪生成单目3D检测伪标签,无需额外传感器
PLOT: Pseudo-Labeling via Object Tracking for Monocular 3D Object Detection
- 通过物体与背景轨迹估计相机运动,实现无姿态假设的物体关联
- 利用帧间点对应融合伪激光雷达,生成抗遮挡和视角变化的统一形状
- 构建全局物体记忆保持身份一致,适用于真实复杂场景
单目3D目标检测对自动驾驶、机器人和监控等领域的可扩展感知至关重要,但受限于稀少的3D标注及单图像几何固有的模糊性。现有方法常依赖强几何假设或精心筛选的数据集,难以推广至真实场景。本文提出PLOT(Pseudo-Labeling via Object Tracking),一种从单目视频中生成3D标注的框架,无需辅助传感器或模型重训练。PLOT通过追踪物体与背景轨迹,估计相机运动并实现姿态未知条件下的物体关联。这些轨迹提供点对应关系,用于对齐帧间伪激光雷达,并通过简单优化融合成统一且鲁棒的物体形状,抵御遮挡与视角变化。考虑到时序一致性是可靠形状融合与视频感知的基本要求,设计全局物体记忆以跨帧保持一致的身份。PLOT在M3OD视频基准和真实场景视频上均实现高质量标注与强泛化能力,证明其在多样且无约束领域中的有效性。
原文摘要 · Abstract (English)
Monocular 3D object detection is crucial for scalable perception across fields like autonomous driving, robotics, and surveillance. However, progress is hindered by limited 3D annotations and the inherent ambiguity of single-image geometry. Existing methods often rely on strong geometric assumptions or carefully curated datasets, which limit their applicability to real-world scenarios. In this paper, we present PLOT (Pseudo-Labeling via Object Tracking), a framework that generates 3D annotations from monocular videos without auxiliary sensors or model retraining. PLOT tracks object and background trajectories to estimate camera motion and perform object association in pose-unknown settings. These trajectories provide point correspondences that align frame-wise pseudo-LiDARs, which are then fused via simple optimization into a unified object shape robust to occlusion and viewpoint shifts. Recognizing temporal coherence as a fundamental requirement for reliable shape fusion and video perception, we design a global object memory that preserves consistent object identities across frames. PLOT achieves robust annotation quality and strong generalization on both M3OD video benchmarks and in-the-wild videos, proving its effectiveness across diverse and unconstrained domains. Project page: https://plot-eccv.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。