arXiv:2506.05199cs.CV2025-06被引 4

统一检测与定位的3D视觉定位框架,性能显著提升。

DEGround: An Effective Baseline for Ego-centric 3D Visual Grounding with a Homogeneous Framework

  • 用共享查询实现检测与定位的统一表示
  • 在EmbodiedScan上整体精度领先7.52%
  • 适合需要精准空间文本对齐的研究者

具身智能中的核心任务是视角中心的3D视觉定位。现有方法多采用两阶段异构流程,将检测器与独立定位模型结合,因解码器和边界框头不兼容,难以传递对象级先验信息,且分步训练导致冗余优化。为此,我们提出DEGround,一种以对象级共享为核心的统一框架。该框架使用一组查询作为检测与定位的共同对象表示,通过共享Transformer和边界框头进行解码。在此统一架构基础上,引入两个任务特异性模块:区域激活定位模块通过突出指令相关区域增强空间-文本对齐;查询级调制模块在初始化时应用句条件仿射调制生成指令感知查询。大量实验表明,DEGround在多个基准上达到最佳性能,尤其在EmbodiedScan数据集上整体精度相比之前方法显著提升7.52%。

原文摘要 · Abstract (English)

A core task in embodied intelligence is ego-centric 3D visual grounding. Existing methods typically adopt two-stage, heterogeneous pipelines that pair a detector with a separate grounding model. Incompatible decoders and box heads hinder the transfer of object-level priors, and the split training causes redundant re-optimization. To overcome these limitations, we present DEGround, a straight, elegant, and effective framework that centers on object-level sharing over detection and grounding. It employs a set of queries that serves as the common object representation for both detection and grounding, which is decoded by a shared transformer and bounding box head. Building on this homogeneous framework, we further introduce two task-specific plug-in modules to enhance fine-grained instruction grounding. The Regional Activation Grounding module improves spatial-textual alignment by highlighting instruction-relevant regions, while the Query-wise Modulation module applies sentence-conditioned affine modulation to generate instruction-aware queries at initialization. Extensive experiments demonstrate that DEGround achieves the best performance on multiple benchmarks. Remarkably, it significantly outperforms previous methods by 7.52% at overall precision on the EmbodiedScan dataset.

3D视觉定位统一框架具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。