用课程式强化学习提升自动驾驶目标检测的鲁棒性
Robust Object Detection for Autonomous Driving via Curriculum-Guided Group Relative Policy Optimization
- 设计课程调度与难度过滤机制,优化稀疏奖励下的训练过程
- 在多个自动驾驶基准上检测精度显著提升,对复杂样本适应更强
- 适合关注多模态感知与强化学习融合的开发者与研究者
多模态大语言模型在视觉-语言推理中表现优异,但在需要精确定位和鲁棒性的结构化感知任务中常遇瓶颈。本文提出一种强化学习框架,通过引入课程引导的数据调度与难度感知过滤,增强群组相对策略优化(GRPO)的稳定性,使其能逐步适应复杂样本。在自动驾驶基准测试中,该方法显著提升了检测准确率与鲁棒性。消融实验验证了奖励设计、KL正则化及课程节奏对收敛稳定性和泛化能力的重要性。研究结果表明,结合结构化数据课程的强化驱动优化,是实现可扩展、可解释多模态检测的有效路径。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) excel in vision-language reasoning but often struggle with structured perception tasks requiring precise localization and robustness. We propose a reinforcement learning framework that augments Group Relative Policy Optimization (GRPO) with curriculum-based data scheduling and difficulty-aware filtering. This approach stabilizes optimization under sparse, noisy rewards and enables progressive adaptation to complex samples. Evaluations on autonomous driving benchmarks demonstrate substantial improvements in detection accuracy and robustness. Ablation studies confirm the importance of reward design, KL regularization, and curriculum pacing for convergence stability and generalization. Our findings highlight reinforcement-driven optimization with structured data curricula as a scalable path toward robust and interpretable multimodal detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。