用预训练3D模型实现跨类别点云追踪,无需为每类物体单独建模。
TrackAny3D: Transferring Pretrained 3D Models for Category-unified 3D Point Cloud Tracking
- 通过轻量适配器和几何专家混合架构,迁移预训练3D模型
- 在三个基准上达到新最好性能,支持跨类别泛化
- 适合需要通用3D追踪能力的研究与工程应用
基于3D LiDAR的单目标追踪(SOT)依赖稀疏不规则的点云,受物体类别间尺度、运动模式和结构复杂度差异影响,现有类别特定方法虽精度高但难以实用,需为每类物体训练独立模型且泛化能力弱。为此,我们提出TrackAny3D,首个面向类别无关3D SOT的预训练模型迁移框架。首先引入参数高效适配器,弥合预训练与追踪任务差距并保留几何先验;其次设计几何专家混合(MoGE)架构,根据物体几何特征自适应激活专用子网络;此外,设计时序上下文优化策略,通过可学习时序标记与动态掩码加权模块,有效传播历史信息并缓解时序漂移。在三个常用基准上的实验表明,TrackAny3D在类别无关3D SOT上建立新最优性能,展现出强大泛化能力与竞争力。希望本工作能启发社区关注统一模型的重要性,并推动大规模预训练模型在该领域的应用拓展。
原文摘要 · Abstract (English)
3D LiDAR-based single object tracking (SOT) relies on sparse and irregular point clouds, posing challenges from geometric variations in scale, motion patterns, and structural complexity across object categories. Current category-specific approaches achieve good accuracy but are impractical for real-world use, requiring separate models for each category and showing limited generalization. To tackle these issues, we propose TrackAny3D, the first framework to transfer large-scale pretrained 3D models for category-agnostic 3D SOT. We first integrate parameter-efficient adapters to bridge the gap between pretraining and tracking tasks while preserving geometric priors. Then, we introduce a Mixture-of-Geometry-Experts (MoGE) architecture that adaptively activates specialized subnetworks based on distinct geometric characteristics. Additionally, we design a temporal context optimization strategy that incorporates learnable temporal tokens and a dynamic mask weighting module to propagate historical information and mitigate temporal drift. Experiments on three commonly-used benchmarks show that TrackAny3D establishes new state-of-the-art performance on category-agnostic 3D SOT, demonstrating strong generalization and competitiveness. We hope this work will enlighten the community on the importance of unified models and further expand the use of large-scale pretrained models in this field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。