用形状和语言联合检索3D物体,提升遮挡环境下的识别准确率。
SAMURAI: Shape-Aware Multimodal Retrieval for 3D Object Identification
- 融合视觉语义与形状轮廓,通过二值轮廓重排序增强检索
- 在ROOMELSA测试集上表现优于基线,关键依赖形状先验
- 适合做开放世界3D物体识别的多模态系统研发
仅凭遮挡的2D图像和自然语言描述,在复杂室内环境中检索3D物体面临巨大挑战。ROOMELSA挑战赛限制了完整3D场景上下文的访问,使物体外观、几何结构和语义推理变得困难。这些问题因视角扭曲、纹理缺失的遮挡区域、模糊的语言提示和噪声分割掩码而加剧。为此,我们提出SAMURAI:面向3D物体识别的形状感知多模态检索方法。SAMURAI结合基于CLIP的语义匹配与由遮挡区域二值轮廓导出的形状引导重排序,并采用鲁棒的多数投票策略。专用预处理流程通过提取最大连通域并去除背景噪声来提升掩码质量。该混合检索框架同时利用语言与形状线索,在ROOMELSA私有测试集上取得竞争力表现。结果表明,将形状先验与语言理解相结合对鲁棒的开放世界3D物体检索至关重要。
原文摘要 · Abstract (English)
Retrieving 3D objects in complex indoor environments using only a masked 2D image and a natural language description presents significant challenges. The ROOMELSA challenge limits access to full 3D scene context, complicating reasoning about object appearance, geometry, and semantics. These challenges are intensified by distorted viewpoints, textureless masked regions, ambiguous language prompts, and noisy segmentation masks. To address this, we propose SAMURAI: Shape-Aware Multimodal Retrieval for 3D Object Identification. SAMURAI integrates CLIP-based semantic matching with shape-guided re-ranking derived from binary silhouettes of masked regions, alongside a robust majority voting strategy. A dedicated preprocessing pipeline enhances mask quality by extracting the largest connected component and removing background noise. Our hybrid retrieval framework leverages both language and shape cues, achieving competitive performance on the ROOMELSA private test set. These results highlight the importance of combining shape priors with language understanding for robust open-world 3D object retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。