arXiv:2511.18151cs.DCcs.AR2025-11被引 4

AVERY让无人机在灾害现场智能分配视觉语言模型任务,兼顾实时与精准。

AVERY: Intent-Driven Adaptive VLM Split Computing via Embodied Self-Awareness for Efficient Disaster Response Systems

  • 按任务意图分两条流:实时感知用低分辨率流,深度分析用高保真流
  • 动态调整压缩模型,在网络波动下仍保持93.98%能耗降低和近似最高精度
  • 适合需要实时响应与精准判断的灾害救援场景

灾害救援中的无人机需具备复杂可查询的智能,但传统本地卷积神经网络难以满足。尽管视觉语言模型(VLMs)能提供语义推理能力,其高资源需求使设备端部署不可行,而直接云端卸载在灾害区低带宽、不稳定网络下也失效。本文提出AVERY,一种基于任务意图的自适应分层计算框架,用于资源受限平台上的高效VLM部署。AVERY的核心思想是将操作员意图作为系统首要目标:广域态势监控与空间精确定位的调查任务,需不同语义产出、延迟目标和资源分配。为此,AVERY突破传统按深度拆分的模式,采用受认知启发的双流结构:高频低分辨率的上下文流用于实时感知,低频高保真的洞察流用于深度分析。该设计支持分层拆分策略:先按功能分离,再在需要洞察流时跨边缘与云端深度划分计算。一个轻量级自感知机载控制器实时监测网络状态与操作意图,动态选择预训练压缩模型,实现运行时精度与吞吐率的权衡。在边缘-云环境下,基于LISA-7B模型的评估显示,AVERY相比原始图像压缩提升11.2%准确率,相比全边缘执行降低93.98%能耗,动态适配期间平均准确率仅比静态高精度基线低0.75%。总体而言,AVERY提升了任务效率,实现在动态灾害环境中的实时可查询智能。

原文摘要 · Abstract (English)

Unmanned Aerial Vehicles (UAVs) in disaster response require complex, queryable intelligence that onboard CNNs cannot provide. While Vision-Language Models (VLMs) offer this semantic reasoning, their high resource demands make on-device deployment infeasible, and naive cloud offloading fails under the low-bandwidth, unstable networks endemic to disaster zones. We present AVERY, an intent-driven adaptive split computing framework for efficient VLM deployment on resource-constrained platforms. AVERY is motivated by the observation that operator intent must be treated as a first-class system objective, since missions such as broad situational monitoring and precise, spatially grounded investigation require different semantic products, latency targets, and resource allocations. To reflect this, AVERY advances split computing beyond traditional depth-wise partitioning through a functional, cognitive-inspired dual-stream split: a high-frequency, low-resolution Context stream for real-time awareness, and a low-frequency, high-fidelity Insight stream for deep analysis. This design enables a hierarchical split strategy: computation is first separated by function, then partitioned depth-wise across edge and cloud when the Insight stream is required. A lightweight, self-aware onboard controller monitors network conditions and operator intent to select from pre-trained compression models, navigating the accuracy-throughput trade-off at runtime. Evaluated using LISA-7B in an edge-cloud setting under fluctuating network conditions, AVERY achieves 11.2% higher accuracy than raw image compression, 93.98% lower energy consumption than full-edge execution, and average accuracy within 0.75% of the static High-Accuracy baseline during dynamic adaptation. Overall, AVERY enhances mission efficiency and enables real-time, queryable intelligence in dynamic disaster environments.

视觉语言模型无人机灾害救援边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。