arXiv:2504.10048cs.CV2025-04被引 1

通过分层对比孪生架构,提升复杂指令下多物体3D定位精度

Multi-Object Grounding via Hierarchical Contrastive Siamese Transformers

  • 分层处理逐步精炼物体定位,结合对比孪生结构增强语义理解
  • 在复杂多物体定位基准上性能超越此前最优方法9.5%
  • 适合需要精准理解复杂语言指令的3D场景理解任务

3D场景中的多物体定位需根据自然语言输入定位多个物体。现有工作多聚焦单物体定位,而真实场景常需同时定位多个物体。为此,我们提出分层对比孪生变压器(H-COST),采用分层处理策略逐步优化物体定位,增强对复杂语言指令的理解。此外,引入对比孪生变压器框架:两个结构相同的网络协同工作,一个辅助网络基于真值标签提取鲁棒的物体关系,引导参考网络(处理分割点云数据);该对比机制强化模型语义理解能力,显著提升处理复杂点云数据的表现。本方法在具有挑战性的多物体定位基准上,性能优于此前最先进方法9.5%。

原文摘要 · Abstract (English)

Multi-object grounding in 3D scenes involves localizing multiple objects based on natural language input. While previous work has primarily focused on single-object grounding, real-world scenarios often demand the localization of several objects. To tackle this challenge, we propose Hierarchical Contrastive Siamese Transformers (H-COST), which employs a Hierarchical Processing strategy to progressively refine object localization, enhancing the understanding of complex language instructions. Additionally, we introduce a Contrastive Siamese Transformer framework, where two networks with the identical structure are used: one auxiliary network processes robust object relations from ground-truth labels to guide and enhance the second network, the reference network, which operates on segmented point-cloud data. This contrastive mechanism strengthens the model' s semantic understanding and significantly enhances its ability to process complex point-cloud data. Our approach outperforms previous state-of-the-art methods by 9.5% on challenging multi-object grounding benchmarks.

3D定位多物体对比学习点云

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。