提出D-LISA模型,实现更精准的多物体3D定位。
Multi-Object 3D Grounding with Dynamic Modules and Language-Informed Spatial Attention
- 动态视觉模块生成可学习的框候选
- 语言引导空间注意力提升定位准确率12.8%
- 适合机器人、人机交互等三维理解场景
多物体3D定位任务旨在根据查询短语从点云中定位3D边界框,具有重要应用价值。本文提出D-LISA,一种两阶段方法,包含三项创新:1)动态视觉模块,支持可变且可学习的框候选数量;2)动态相机定位机制,为每个候选框提取特征;3)语言引导的空间注意力模块,优化候选框推理。实验表明,该方法在多物体3D定位上相较当前最优方法提升12.8%(绝对值),在单物体任务中也表现优异。
原文摘要 · Abstract (English)
Multi-object 3D Grounding involves locating 3D boxes based on a given query phrase from a point cloud. It is a challenging and significant task with numerous applications in visual understanding, human-computer interaction, and robotics. To tackle this challenge, we introduce D-LISA, a two-stage approach incorporating three innovations. First, a dynamic vision module that enables a variable and learnable number of box proposals. Second, a dynamic camera positioning that extracts features for each proposal. Third, a language-informed spatial attention module that better reasons over the proposals to output the final prediction. Empirically, experiments show that our method outperforms the state-of-the-art methods on multi-object 3D grounding by 12.8% (absolute) and is competitive in single-object 3D grounding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。