arXiv:2602.13324cs.CVcs.AI2026-02被引 1

零样本框架让边缘机器人自主识别目标并推理战术,准确率超97%。

Synthesizing the Kill Chain: A Zero-Shot Framework for Target Verification and Tactical Reasoning on the Edge

  • 用轻量检测+小规模视觉语言模型分层处理,实现边缘端零样本推理。
  • 在合成视频上实现97.5%损伤评估准确率,目标部署正确率达100%。
  • 揭示大模型在感知与推理中的不同失效模式,适合军事边缘智能应用。

在动态军事环境中部署自主边缘机器人受限于领域特定训练数据稀缺和边缘硬件计算能力不足。本文提出一种分层零样本框架,将轻量级目标检测与小型视觉语言模型(来自Qwen和Gemma系列,4B-12B参数)级联使用。Grounding DINO作为高召回、文本提示驱动的区域提议器,将高置信度帧传递给边缘分类VLM进行语义验证。我们在55段来自Battlefield 6的高保真合成视频上评估该流程,涵盖三项任务:误报过滤(最高100%准确率)、损伤评估(最高97.5%)和细粒度车辆分类(55%-90%)。进一步扩展为代理式Scout-Commander工作流,实现100%正确资产部署,推理评分达9.8/10(GPT-4o评分),延迟低于75秒。提出的“受控输入”方法解耦感知与推理,揭示不同故障特征:Gemma3-12B在战术逻辑上表现优异但视觉感知失败;Gemma3-4B即使输入准确也出现推理崩溃。结果验证了分层零样本架构在边缘自主中的有效性,并为安全关键场景下VLM适用性提供了诊断框架。

原文摘要 · Abstract (English)

Deploying autonomous edge robotics in dynamic military environments is constrained by both scarce domain-specific training data and the computational limits of edge hardware. This paper introduces a hierarchical, zero-shot framework that cascades lightweight object detection with compact Vision-Language Models (VLMs) from the Qwen and Gemma families (4B-12B parameters). Grounding DINO serves as a high-recall, text-promptable region proposer, and frames with high detection confidence are passed to edge-class VLMs for semantic verification. We evaluate this pipeline on 55 high-fidelity synthetic videos from Battlefield 6 across three tasks: false-positive filtering (up to 100% accuracy), damage assessment (up to 97.5%), and fine-grained vehicle classification (55-90%). We further extend the pipeline into an agentic Scout-Commander workflow, achieving 100% correct asset deployment and a 9.8/10 reasoning score (graded by GPT-4o) with sub-75-second latency. A novel "Controlled Input" methodology decouples perception from reasoning, revealing distinct failure phenotypes: Gemma3-12B excels at tactical logic but fails in visual perception, while Gemma3-4B exhibits reasoning collapse even with accurate inputs. These findings validate hierarchical zero-shot architectures for edge autonomy and provide a diagnostic framework for certifying VLM suitability in safety-critical applications.

边缘智能零样本战术推理视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。