arXiv:2608.01495cs.CV2026-08

2D检测模型竟隐含3D物体信息,无需3D训练也能恢复深度和位置。

Probing the 3D Object-Level Understanding of Pre-Trained Detection Transformers

论文配图:Probing the 3D Object-Level Understanding of Pre-Trained Detection Transformers
图 1 · 摘自论文原文
  • 用线性与非线性探针从对象嵌入中提取3D属性
  • 多种模型在无3D监督下仍能准确恢复物体深度和相对位置
  • 适用于对3D理解感兴趣的视觉模型研究者

检测变压器模型(如DETR及其扩展)通过输出一组对象级嵌入,可同时解码出2D边界框和类别分布。本文探究预训练的2D检测变压器对物体3D属性的理解能力,重点考察物体相对于相机的深度及3D位置能否通过线性与非线性探针从对象级嵌入中恢复。实验覆盖多种检测变压器模型,结果显示:尽管预训练过程中完全缺乏3D监督,这些2D模型仍表现出惊人且此前未知的能力,能够有效表征物体的3D属性。

原文摘要 · Abstract (English)

Detection transformer models, including DETR and its extensions, learn to output a set of object-level embeddings that can be simultaneously decoded into 2D bounding boxes and class distributions. In this paper, we investigate what pre-trained 2D detection transformers understand about the 3D properties of objects. Specifically, we investigate the extent to which properties including the depth of objects from the camera and the 3D location of objects relative to the camera can be recovered from object-level embeddings using linear and non-linear probes. Across a range of detection transformer models, our results show a surprisingly strong and previously unknown ability of 2D DETR models to represent useful information about the 3D properties of objects, despite the complete lack of 3D supervision during model pre-training.

目标检测3D理解视觉模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。