arXiv:2607.18874cs.LGcs.CY2026-07

用强化学习优化无人机在风扰环境下的配送与感知协同调度。

Reinforcement Learning for Delivery Drone-Based Participatory Sensing in Dynamic Environments

论文配图:Reinforcement Learning for Delivery Drone-Based Participatory Sensing in Dynamic Environments
图 1 · 摘自论文原文
  • 分时序双层强化学习,宏观任务调度与微观速度控制分离决策。
  • 上海、杭州实测系统收益分别提升46.6%和20.1%。
  • 适合城市级无人机群动态感知场景,尤其关注环境扰动的系统设计。

利用无人机进行城市感知已成为监测空气质量、噪声水平等城市状态的高效方式,通过灵活的空中众包实现。然而,现有基于无人机的感知方法忽略了风力等环境干扰对飞行速度和能效的显著影响。因此,在动态环境下直接应用现有方法于配送与感知融合场景时面临两大挑战:(1)随着机群规模扩大,可扩展性瓶颈凸显;(2)宏观任务调度与微观速度控制存在多时间尺度决策异质性。为此,我们提出SensUAV问题并设计两时间尺度强化学习框架(TSRL)。TSRL将决策分为两个协同层级:宏观层面,任务嵌入式感知调度器通过独立编码任务特征并依次评估无人机适配度完成任务选择,保障可扩展性;微观层面,风感知速度控制器学习精细速度调度以适应动态环境变化。在真实数据集上的大量实验表明,TSRL显著优于基线方法,在杭州和上海分别实现平均系统收益提升20.1%和46.6%。

原文摘要 · Abstract (English)

Using Unmanned Aerial Vehicle (UAV) for urban sensing has emerged as a powerful paradigm to monitor the status of the city, e.g., air quality and noise levels, through agile aerial crowdsourcing. Despite this potential, existing UAV-based sensing approaches overlook environmental disturbances like wind that drastically impact drone velocity and energy efficiency. Consequently, directly applying existing methods to this joint delivery and sensing paradigm in dynamic environments faces two severe challenges: (1) scalability bottlenecks as fleet sizes expand; and (2) multi-timescale decision heterogeneity between macro task dispatching and micro velocity control. To tackle these, we formalize the problem as SensUAV and propose a Two TimeScale Reinforcement Learning framework (TSRL). Specifically, TSRL separates decision-making into two cooperative layers. At the macro level, a task-embedding sensing dispatcher handles scalability by separately encoding distinct task features and sequentially evaluating UAV suitability before task selection. At the micro level, a wind-aware velocity controller learns fine-grained velocity scheduling to adapt to dynamic environmental variations. Extensive experiments on real-world datasets demonstrate that TSRL significantly outperforms baselines, achieving average system profit improvements of 20.1% in Hangzhou and 46.6% in Shanghai.

无人机强化学习城市感知动态调度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。