用场景提示池提升3D目标检测的迁移能力
SOP^2: Transfer Learning with Scene-Oriented Prompt Pool on 3D Object Detection
- 设计场景感知提示池,动态适配不同环境下的检测任务
- 在Waymo数据集上训练的模型可有效迁移到其他场景
- 为3D视觉领域提示工程提供新思路,适合做迁移学习研究者
随着GPT-3等大语言模型的发展,其强大的泛化能力通过微调和提示调优等迁移学习技术,可在少量参数调整下适应多种下游任务,这一方法在自然语言处理中已广泛应用。本文探索提示调优在3D目标检测中的有效性,研究在大规模Waymo数据集上训练的模型能否作为基础模型,适配其他3D检测场景。论文系统考察了提示词与提示生成器的影响,并进一步提出场景导向提示池(SOP²)。实验证明提示池在3D目标检测中具有显著效果,旨在激发未来研究者对提示机制在3D领域潜力的深入探索。
原文摘要 · Abstract (English)
With the rise of Large Language Models (LLMs) such as GPT-3, these models exhibit strong generalization capabilities. Through transfer learning techniques such as fine-tuning and prompt tuning, they can be adapted to various downstream tasks with minimal parameter adjustments. This approach is particularly common in the field of Natural Language Processing (NLP). This paper aims to explore the effectiveness of common prompt tuning methods in 3D object detection. We investigate whether a model trained on the large-scale Waymo dataset can serve as a foundation model and adapt to other scenarios within the 3D object detection field. This paper sequentially examines the impact of prompt tokens and prompt generators, and further proposes a Scene-Oriented Prompt Pool (\textbf{SOP$^2$}). We demonstrate the effectiveness of prompt pools in 3D object detection, with the goal of inspiring future researchers to delve deeper into the potential of prompts in the 3D field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。