构建新基准,评估模型在复杂场景中根据描述精准找3D物体的能力。
SHREC 2025: Retrieval of Optimal Objects for Multi-modal Enhanced Language and Spatial Assistance (ROOMELSA)
- 基于全景图定位目标区域,结合自然语言检索3D模型。
- 含超4.4万条查询,仅1个模型在近全测试集上首名命中。
- 适合研究多模态理解与真实场景3D识别的学者参考。
当前3D检索系统多针对简单、受控场景,如从裁剪图像或简短描述中识别物体。然而,现实场景更复杂,常需基于模糊、自由描述在杂乱场景中识别物体。为此,我们提出ROOMELSA,一个新基准,用于评估系统理解自然语言的能力。具体而言,ROOMELSA聚焦全景房间图像中的特定区域,从大型数据库中准确检索对应3D模型。此外,ROOMELSA包含超过1,600个公寓场景、近5,200个房间和超过44,000条目标查询。实证表明,尽管粗粒度物体检索已基本解决,但仅有1个顶尖模型在近所有测试案例中将正确匹配排在首位。值得注意的是,一个轻量级CLIP-based模型表现良好,但仍难以处理材质、部件结构及上下文线索的细微差异,导致偶发错误。这些发现凸显了视觉与语言理解紧密集成的重要性。通过弥合场景级定位与细粒度3D检索之间的差距,ROOMELSA为推进鲁棒的现实世界3D识别系统设立了新基准。
原文摘要 · Abstract (English)
Recent 3D retrieval systems are typically designed for simple, controlled scenarios, such as identifying an object from a cropped image or a brief description. However, real-world scenarios are more complex, often requiring the recognition of an object in a cluttered scene based on a vague, free-form description. To this end, we present ROOMELSA, a new benchmark designed to evaluate a system's ability to interpret natural language. Specifically, ROOMELSA attends to a specific region within a panoramic room image and accurately retrieves the corresponding 3D model from a large database. In addition, ROOMELSA includes over 1,600 apartment scenes, nearly 5,200 rooms, and more than 44,000 targeted queries. Empirically, while coarse object retrieval is largely solved, only one top-performing model consistently ranked the correct match first across nearly all test cases. Notably, a lightweight CLIP-based model also performed well, although it struggled with subtle variations in materials, part structures, and contextual cues, resulting in occasional errors. These findings highlight the importance of tightly integrating visual and language understanding. By bridging the gap between scene-level grounding and fine-grained 3D retrieval, ROOMELSA establishes a new benchmark for advancing robust, real-world 3D recognition systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。