arXiv:2601.14092cs.LGcs.NI2026-01

用注意力多目标强化学习优化无人机采数据与耗能的平衡

Optimizing Energy and Data Collection in UAV-aided IoT Networks using Attention-based Multi-Objective Reinforcement Learning

  • 基于注意力机制的多目标强化学习,动态权衡数据采集与能耗
  • 无需先验信道知识,单模型适应不同场景与偏好设置
  • 在未见场景中表现更优,样本效率与泛化能力显著提升

由于具备适应性和机动性,无人机(UAV)正日益成为无线网络服务的关键,尤其在数据采集任务中。当前基于人工智能的方法虽被广泛研究,但多数算法受限于训练数据不足,在高度动态环境中性能不佳,且常忽略任务的多目标本质。为此,本文提出一种基于注意力的多目标强化学习(MORL)架构,显式处理城市环境中数据收集与能量消耗之间的权衡,无需事先了解无线信道状况。所提方法构建单一模型,可无需微调或重训练即适应不同权衡偏好和动态环境参数。大量仿真表明,该方法在性能、模型紧凑性、样本效率及对未见场景的泛化能力上均显著优于现有强化学习方案。

原文摘要 · Abstract (English)

Due to their adaptability and mobility, Unmanned Aerial Vehicles (UAVs) are becoming increasingly essential for wireless network services, particularly for data harvesting tasks. In this context, Artificial Intelligence (AI)-based approaches have gained significant attention for addressing UAV path planning tasks in large and complex environments, bridging the gap with real-world deployments. However, many existing algorithms suffer from limited training data, which hampers their performance in highly dynamic environments. Moreover, they often overlook the inherently multi-objective nature of the task, treating it in an overly simplistic manner. To address these limitations, we propose an attention-based Multi-Objective Reinforcement Learning (MORL) architecture that explicitly handles the trade-off between data collection and energy consumption in urban environments, even without prior knowledge of wireless channel conditions. Our method develops a single model capable of adapting to varying trade-off preferences and dynamic scenario parameters without the need for fine-tuning or retraining. Extensive simulations show that our approach achieves substantial improvements in performance, model compactness, sample efficiency, and most importantly, generalization to previously unseen scenarios, outperforming existing RL solutions.

无人机网络强化学习多目标优化物联网

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。