arXiv:2511.18814cs.CV2025-11

端到端实现流式视频中4D物体的稳定检测,提升时空一致性。

DetAny4D: Detect Anything 4D Temporally in a Streaming RGB Video

  • 基于多模态融合与几何感知时序解码器,直接从序列输入预测3D框。
  • 在28万+序列数据上实现高精度检测,显著降低运动抖动和不一致问题。
  • 适合自动驾驶、机器人视觉等需要连续感知的实时应用。

可靠的4D物体检测(即流式视频中的3D检测)对理解真实世界至关重要。现有开放集4D检测方法通常逐帧预测且缺乏时序一致性建模,或依赖复杂多阶段流程,易产生误差传播。该领域进展受限于缺乏大规模连续高质量3D边界框标注数据集。为此,我们首次提出DA4D,一个包含超过28万条序列的大规模4D检测数据集,涵盖多种场景下的高精度边界框标注。在此基础上,我们提出DetAny4D,一种开放集端到端框架,直接从序列输入预测3D边界框。DetAny4D融合预训练基础模型的多模态特征,并设计几何感知时空解码器以有效捕捉空间与时间动态。同时采用多任务学习架构及专用训练策略,保持不同长度序列间的全局一致性。大量实验表明,DetAny4D在检测精度上表现优异,且显著提升时序稳定性,有效解决4D检测中长期存在的抖动与不一致问题。数据与代码将在论文接收后公开。

原文摘要 · Abstract (English)

Reliable 4D object detection, which refers to 3D object detection in streaming video, is crucial for perceiving and understanding the real world. Existing open-set 4D object detection methods typically make predictions on a frame-by-frame basis without modeling temporal consistency, or rely on complex multi-stage pipelines that are prone to error propagation across cascaded stages. Progress in this area has been hindered by the lack of large-scale datasets that capture continuous reliable 3D bounding box (b-box) annotations. To overcome these challenges, we first introduce DA4D, a large-scale 4D detection dataset containing over 280k sequences with high-quality b-box annotations collected under diverse conditions. Building on DA4D, we propose DetAny4D, an open-set end-to-end framework that predicts 3D b-boxes directly from sequential inputs. DetAny4D fuses multi-modal features from pre-trained foundational models and designs a geometry-aware spatiotemporal decoder to effectively capture both spatial and temporal dynamics. Furthermore, it adopts a multi-task learning architecture coupled with a dedicated training strategy to maintain global consistency across sequences of varying lengths. Extensive experiments show that DetAny4D achieves competitive detection accuracy and significantly improves temporal stability, effectively addressing long-standing issues of jitter and inconsistency in 4D object detection. Data and code will be released upon acceptance.

4D检测时序建模自动驾驶多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。