提出高效多智能体感知预测训练框架,自动优化多任务学习
TurboTrain: Towards Efficient and Balanced Multi-Task Learning for Multi-Agent Perception and Prediction
- 基于掩码重建的时空预训练+梯度冲突抑制的平衡策略
- 在V2XPnP-Seq数据集上提升现有模型性能,训练时间显著减少
- 适合自动驾驶多智能体系统研发者,尤其关注训练效率与精度
端到端训练多智能体系统能有效提升多任务表现,但训练过程仍具挑战性,需大量人工设计与监控。本文提出TurboTrain,一种新型高效的多智能体感知与预测训练框架。其包含两个核心组件:基于掩码重建学习的多智能体时空预训练方案,以及基于梯度冲突抑制的平衡多任务学习策略。该框架简化了训练流程,无需手动设计复杂多阶段训练管道,大幅缩短训练时间并提升性能。我们在真实世界协同驾驶数据集V2XPnP-Seq上评估了TurboTrain,结果表明其进一步提升了现有先进多智能体感知与预测模型的表现。实验显示,预训练能有效捕捉多智能体时空特征,并显著促进下游任务。此外,所提平衡多任务学习策略增强了检测与预测能力。
原文摘要 · Abstract (English)
End-to-end training of multi-agent systems offers significant advantages in improving multi-task performance. However, training such models remains challenging and requires extensive manual design and monitoring. In this work, we introduce TurboTrain, a novel and efficient training framework for multi-agent perception and prediction. TurboTrain comprises two key components: a multi-agent spatiotemporal pretraining scheme based on masked reconstruction learning and a balanced multi-task learning strategy based on gradient conflict suppression. By streamlining the training process, our framework eliminates the need for manually designing and tuning complex multi-stage training pipelines, substantially reducing training time and improving performance. We evaluate TurboTrain on a real-world cooperative driving dataset, V2XPnP-Seq, and demonstrate that it further improves the performance of state-of-the-art multi-agent perception and prediction models. Our results highlight that pretraining effectively captures spatiotemporal multi-agent features and significantly benefits downstream tasks. Moreover, the proposed balanced multi-task learning strategy enhances detection and prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。