用视觉语言模型让无人车团队像指挥官一样自主决策
Tactical Decision for Multi-UGV Confrontation with a Vision-Language Model-Based Commander
- 结合视觉语言模型与轻量大模型,实现感知到决策的统一推理
- 在仿真中击败基线模型,胜率超80%
- 决策过程可解释,适合需要透明性的军事智能系统
在多无人地面车辆对抗中,从态势感知自主生成战术决策仍是重大挑战。传统基于规则的方法在复杂多变的战场环境下易失效,而现有强化学习方法因缺乏可解释性,主要关注动作控制而非战略决策。本文提出一种基于视觉语言模型的指挥官框架,整合视觉语言模型进行场景理解,以及轻量级大语言模型进行战略推理,实现在共享语义空间内的统一感知与决策,具备强适应性与可解释性。与基于规则搜索和强化学习的方法不同,双模块结合构建了完整的决策链,模拟人类指挥官的认知过程。仿真与消融实验表明,该方法相比基线模型胜率超过80%。
原文摘要 · Abstract (English)
In multiple unmanned ground vehicle confrontations, autonomously evolving multi-agent tactical decisions from situational awareness remain a significant challenge. Traditional handcraft rule-based methods become vulnerable in the complicated and transient battlefield environment, and current reinforcement learning methods mainly focus on action manipulation instead of strategic decisions due to lack of interpretability. Here, we propose a vision-language model-based commander to address the issue of intelligent perception-to-decision reasoning in autonomous confrontations. Our method integrates a vision language model for scene understanding and a lightweight large language model for strategic reasoning, achieving unified perception and decision within a shared semantic space, with strong adaptability and interpretability. Unlike rule-based search and reinforcement learning methods, the combination of the two modules establishes a full-chain process, reflecting the cognitive process of human commanders. Simulation and ablation experiments validate that the proposed approach achieves a win rate of over 80% compared with baseline models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。