提出首个通用单目时序3D检测方法,提升跨数据集泛化能力。
MAGneT-3D: Monocular and Domain-Generalizable Temporal 3D Detection
- 用动态生成的3D候选框替代固定查询,增强环境适应性。
- 零样本跨数据集测试下,NDS指标从12.1%提升至18.6%。
- 适用于自动驾驶中复杂多变场景的鲁棒目标检测。
单目时序3D检测旨在仅通过单目视频实现三维物体检测。基于查询的3D检测器虽统一了检测与跨视角关联,但其可学习查询依赖训练数据的空间分布(如视场范围),在应用于单目视频时泛化能力显著下降。为此,本文提出MAGneT-3D,首个支持域泛化的单目时序3D检测方法。不依赖静态可学习查询,而是提出域鲁棒锚框生成器(DRAG),在推理阶段自适应生成3D候选框;进一步设计时序精修与身份合并策略(TRIM),降低对特定3D候选框的依赖。为支持全面的域泛化评估,构建覆盖nuScenes、Waymo、Lyft、ONCE的跨数据集基准。在零样本域迁移条件下,MAGneT-3D优于所有基线模型,将NDS从12.1%提升至18.6%,同时保持域内检测精度。
原文摘要 · Abstract (English)
Monocular temporal 3D detection aims to detect objects in 3D, given a monocular video. Query-based 3D detectors unify detection and cross-view association, but their learnable queries fit the spatial distribution of the training data (e.g., field-of-view). We show that this issue is especially severe when these models are applied to monocular video, hindering generalization to unseen datasets and environments. To address this limitation, we introduce MAGneT-3D, the first method for domain-generalized monocular temporal 3D object detection. Instead of relying on static learnable queries, we propose a Domain-Robust Anchor Generator (DRAG) approach that adaptively derives 3D proposals during inference. To further enable domain generalization, we propose a Temporal Refinement and Identity Merging (TRIM) strategy, reducing dependence on specific 3D proposals. To enable comprehensive domain-generalization evaluation, we establish a cross-dataset benchmark spanning nuScenes, Waymo, Lyft, and ONCE. Under zero-shot domain shifts, MAGneT-3D outperforms all baselines, improving NDS from 12.1% to 18.6% while also increasing in-domain accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。