构建首个严格对齐时序图像的3D视觉定位基准,提升复杂场景定位精度。
SeqAlign3DVG: A Sequence-Aligned Benchmark and Voxel Reasoning Framework for 3D Visual Grounding

- 基于体素的统一框架,动态聚合多视角证据以抑制噪声。
- 在14,493个序列样本上实现最优定位性能,尤其擅长复杂关系描述。
- 适合研究具身智能、视觉语言导航与多模态推理的学者。
基于图像的3D视觉定位对具身智能体至关重要,但现有基准普遍存在文本-观测对齐松散和忽略时序顺序的问题。本文提出SeqAlign3DVG,一个专注于时序有序且严格观测对齐的图像型3D视觉定位新基准。不同于以往使用无序视角或全局点云的方法,该基准确保所有表达均经人工验证,并严格锚定于提供的RGB观测(单帧或有序观测序列)。数据集包含9,622个单视图样本和14,493个序列样本,涵盖丰富描述、复杂关系及多实例歧义。为应对该基准,我们提出统一的体素化流水线,包含相关性排序体素记忆(ROVM)与渐进式语言-体素融合(PLVF)。ROVM通过保守记忆动态排序并聚合多视角证据以缓解噪声干扰;PLVF执行从粗到细的空间-语言推理,实现精准消歧。所提方法在无深度协议下达到当前最佳性能,显著提升了由复杂关系与外观线索定义目标的定位效果。
原文摘要 · Abstract (English)
Image-based 3D visual grounding is critical for embodied agents, yet existing benchmarks suffer from loose text-observation alignment and neglect temporal ordering. We introduce SeqAlign3DVG, a novel benchmark dedicated to temporally ordered and strictly observation-aligned image-based 3D visual grounding. Unlike prior works using order-agnostic views or global point clouds, SeqAlign3DVG ensures all expressions are human-verified and strictly grounded in the provided RGB observations (single frames or ordered observation sequences). It comprises 9,622 single-view and 14,493 sequence samples featuring rich descriptions, complex relations, and multi-instance ambiguities. To tackle this benchmark, we propose a unified voxel-based pipeline featuring Relevance-Ordered Voxel Memory (ROVM) and Progressive Language-Voxel Fusion (PLVF). ROVM dynamically ranks and aggregates multi-view evidence via a conservative memory to mitigate noisy observations, while PLVF performs coarse-to-fine spatial-linguistic reasoning for precise disambiguation. Our approach achieves state-of-the-art performance under the depth-free protocol, significantly improving localization for targets defined by complex relations and appearance cues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。