arXiv:2412.01539cs.CVcs.RO2024-12被引 18

简化3D场景图设计,提速三倍且不降精度

The Bare Necessities: Designing Simple, Effective Open-Vocabulary Scene Graphs

  • 剔除冗余图像预处理,优化特征融合与选择策略
  • 性能持平顶尖水平,计算量降低三倍
  • 适合追求高效部署的机器人导航研究者

3D开放词汇场景图是具身智能体的有前景地图表示,但现有方法计算成本高昂。本文重新审视以往工作中的关键设计选择,以兼顾效率与性能。提出通用场景图框架,并开展三项研究:图像预处理、特征融合与特征选择。结果表明,常用图像预处理虽提升微弱却使计算量增加三倍(按每物体视角计);跨视角平均特征标签会显著降低性能。通过探索替代特征选择策略,在不增加计算开销的前提下提升性能。基于此,提出一种计算均衡的3D点云分割方法,保留每物体特征,达到与当前最优分类准确率相当,同时实现计算量三倍减少。

原文摘要 · Abstract (English)

3D open-vocabulary scene graph methods are a promising map representation for embodied agents, however many current approaches are computationally expensive. In this paper, we reexamine the critical design choices established in previous works to optimize both efficiency and performance. We propose a general scene graph framework and conduct three studies that focus on image pre-processing, feature fusion, and feature selection. Our findings reveal that commonly used image pre-processing techniques provide minimal performance improvement while tripling computation (on a per object view basis). We also show that averaging feature labels across different views significantly degrades performance. We study alternative feature selection strategies that enhance performance without adding unnecessary computational costs. Based on our findings, we introduce a computationally balanced approach for 3D point cloud segmentation with per-object features. The approach matches state-of-the-art classification accuracy while achieving a threefold reduction in computation.

3D场景图点云分割效率优化具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。