用大模型生成代码解析3D空间关系,实现高效精准的零训练3D视觉定位。
Language-to-Space Programming for Training-Free 3D Visual Grounding
- 用大模型生成代码分析物体间3D空间关系
- 在Nr3D上达52.9%准确率,优于现有零训练方法
- 显著降低推理时间与调用成本,适合实际部署
3D视觉定位因需理解三维空间关系而具挑战性。监督方法虽性能优越,但受限于高质量3D视觉语言数据集稀缺及标注成本高昂。基于大模型/视觉语言模型的零训练方法虽免去大规模训练,却存在推理耗时长、令牌消耗大或准确率不理想的问题。为此,本文提出一种全新的零训练3D视觉定位方法——语言到空间编程(LaSP)。LaSP利用大模型生成代码来分析物体间的3D空间关系,并构建自动评估与优化代码的流水线。实验表明,LaSP在Nr3D基准上达到52.9%的准确率,位居当前最优零训练方法之列,且显著降低推理时间和令牌开销,实现了性能与效率的平衡。
原文摘要 · Abstract (English)
3D visual grounding (3DVG) is challenging due to the need to understand 3D spatial relations. While supervised approaches have achieved superior performance, they are constrained by the scarcity and high annotation costs of 3D vision-language datasets. Training-free approaches based on LLMs/VLMs eliminate the need for large-scale training data, but they either incur prohibitive grounding time and token costs or have unsatisfactory accuracy. To address the challenges, we introduce a novel method for training-free 3D visual grounding, namely Language-to-Space Programming (LaSP). LaSP introduces LLM-generated codes to analyze 3D spatial relations among objects, along with a pipeline that evaluates and optimizes the codes automatically. Experimental results demonstrate that LaSP achieves 52.9% accuracy on the Nr3D benchmark, ranking among the best training-free methods. Moreover, it substantially reduces the grounding time and token costs, offering a balanced trade-off between performance and efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。