arXiv:2603.11447cs.RO2026-03

轻量视觉语言模型通过竞争学习实现高效社交导航

Enhancing Lightweight Vision Language Models through Group Competitive Learning for Socially Compliant Navigation

  • 提出分组竞争学习框架,融合全局语义与分布正则化
  • 3B模型经训练后超越8B基线模型28%性能,F1达0.968
  • 适合需实时推理的机器人部署场景,兼顾精度与效率

社交机器人导航需融合场景语义与人类社交规范。尽管扩大视觉语言模型(VLM)规模通常能提升决策能力,但计算开销大,难用于实时机器人部署。轻量VLM虽推理高效,但在复杂社交环境中推理能力弱。为弥合此差距,本文提出分组竞争学习(GCL),引入分组竞争目标(GCO)协调全局语义与分布正则化,并采用非对称分组优化(AGO)探索模型性能上限。在社交导航基准上评估显示,GCL显著提升性能:Qwen2.5-VL-3B模型在微调后达到F1=0.968,较普通监督微调(SFT)提升40%;而原3B模型落后于8B模型(F1: 0.692 vs. 0.755),经GCL后反超28%。结果表明,GCL是实现实时部署下高精度与高效率的有效方案。

原文摘要 · Abstract (English)

Social robot navigation requires a sophisticated integration of scene semantics and human social norms. Scaling up Vision Language Models (VLMs) generally improves reasoning and decision-making capabilities for socially compliant navigation. However, increased model size incurs substantial computational overhead, limiting suitability for real-time robotic deployment. Conversely, lightweight VLMs enable efficient inference but often exhibit weaker reasoning and decision-making performance in socially complex environments. Achieving both strong reasoning ability and efficiency remains an open challenge. To bridge this gap, we propose Group Competitive Learning (GCL), a strategy designed to amplify the capabilities of lightweight VLMs. Our strategy introduces the Group Competitive Objective (GCO) to harmonize global semantics with distributional regularization, alongside Asymmetric Group Optimization (AGO) to explore the upper limits of model performance. Empirical evaluations on social navigation benchmarks demonstrate that GCL significantly elevates VLM performance. Specifically, GCL enables the Qwen2.5-VL-3B learner model and guide Qwen3-VL-4B to achieve an F1 score of 0.968 and 0.914, representing 40\% and 12\% improvement over vanilla supervised fine-tuning (SFT). Notably, under vanilla SFT, the 3B model initially trails the 8B model (F1: 0.692 vs. 0.755). However, through the GCL, the 3B model outperforms (28\%) the 8B baseline model. These results suggest that GCL provides an effective solution for achieving both high accuracy and computational efficiency in real-world deployment.

视觉语言模型机器人导航轻量化竞争学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。