arXiv:2510.10028cs.LGcs.AI2025-10被引 2

用大模型优化无人机视觉问答,提升低空网络实时性与能效。

Efficient Onboard Vision-Language Inference in UAV-Enabled Low-Altitude Economy Networks via LLM-Enhanced Optimization

  • 分层优化框架:资源分配与轨迹规划协同设计。
  • 任务延迟降低37%,功耗减少41%,满足用户精度要求。
  • 大模型离线助训强化学习,实现实时无额外开销决策。

低空经济网络(LAENets)的快速发展推动了航拍监控、环境感知和语义数据采集等应用。为支持这些场景,搭载视觉语言模型(VLMs)的无人机(UAVs)可实现近端多模态推理。然而,在有限机载资源和动态网络条件下,保障推理准确率与通信效率仍具挑战。本文提出一种融合无人机移动性、用户-无人机通信及机载视觉问答(VQA)流程的系统模型,构建混合整数非凸优化问题,以在满足用户精度约束下最小化任务延迟与功耗。设计分层优化框架:(i)交替分辨率与功率优化(ARPO)算法用于带精度约束的资源分配;(ii)大语言模型增强的强化学习方法(LLaRA)实现自适应无人机轨迹优化。大语言模型在离线阶段辅助优化强化学习奖励函数,不引入实时决策延迟。数值结果表明,该框架在动态LAENet环境下显著提升推理性能与通信效率。

原文摘要 · Abstract (English)

The rapid advancement of Low-Altitude Economy Networks (LAENets) has enabled a variety of applications, including aerial surveillance, environmental sensing, and semantic data collection. To support these scenarios, unmanned aerial vehicles (UAVs) equipped with onboard vision-language models (VLMs) offer a promising solution for real-time multimodal inference. However, ensuring both inference accuracy and communication efficiency remains a significant challenge due to limited onboard resources and dynamic network conditions. In this paper, we first propose a UAV-enabled LAENet system model that jointly captures UAV mobility, user-UAV communication, and the onboard visual question answering (VQA) pipeline. Based on this model, we formulate a mixed-integer non-convex optimization problem to minimize task latency and power consumption under user-specific accuracy constraints. To solve the problem, we design a hierarchical optimization framework composed of two parts: (i) an Alternating Resolution and Power Optimization (ARPO) algorithm for resource allocation under accuracy constraints, and (ii) a Large Language Model-augmented Reinforcement Learning Approach (LLaRA) for adaptive UAV trajectory optimization. The large language model (LLM) serves as an expert in refining reward design of reinforcement learning in an offline fashion, introducing no additional latency in real-time decision-making. Numerical results demonstrate the efficacy of our proposed framework in improving inference performance and communication efficiency under dynamic LAENet conditions.

无人机视觉问答优化大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。