arXiv:2605.02705cs.LGcs.NI2026-05

在信息不全下,用联邦强化学习让手机众包用户高效接任务。

Federated Reinforcement Learning for Efficient Mobile Crowdsensing under Incomplete Information

  • 各手机用户自主学策略,不依赖全局信息
  • 任务完成率、公平性和能耗均优于现有方法
  • 适合能源受限的分布式移动众包场景

移动众包(MCS)利用手机终端(MUs)的传感器完成感知任务,平台发布任务并支付报酬。系统动态变化:任务需求、用户可用性及资源随时间波动。用户需在信息不全的情况下制定参与策略以最大化收入,平台则关注任务完成数量。为应对非因果信息缺失的挑战,提出一种完全去中心化的联邦深度强化学习算法FDRL-PPO。该算法使每个MU基于自身经验、资源和偏好自主学习参与策略,无需全局信息。由于依赖能量采集补能,用户能量水平动态变化,导致可用性波动和学习体验碎片化。所提方法通过联邦学习实现模型共享而非原始数据交换,使用户协同提升模型性能,弥补个体局限。在合成与真实数据集上的评估表明,FDRL-PPO在任务完成率、任务分配公平性、能耗及冲突提案数方面均显著优于基准算法。

原文摘要 · Abstract (English)

Mobile crowdsensing (MCS) is a distributed sensing architecture that utilizes existing sensors on mobile units (MUs) to perform sensing tasks. A mobile crowdsensing platform (MCSP) publishes the sensing tasks and the MUs decide whether to participate in exchange for money. The MCS system is dynamic: the task requirements, the MUs' availability, and their available resources change over time. The MUs aim to find an efficient task participation strategy to maximize their income while the MCSP focuses on maximizing the number of completed tasks. As optimal strategies require perfect non-causal information about the MCS system, which is unavailable in realistic scenarios, the main challenge is to find an efficient task participation strategy for the MUs under incomplete information. To this end, a novel fully decentralized federated deep reinforcement learning algorithm, FDRL-PPO, is proposed. FDRL-PPO enables every MU to learn its own task participation strategy based on its experiences, available resources, and preferences, without relying on perfect non-causal information about the MCS system. To replenish their batteries, the MUs rely on energy harvesting. As a result, their available energy varies over time, leading to varying availability and fragmented learning experiences. To mitigate these challenges, the proposed approach leverages federated learning, enabling MUs to collaboratively improve their models without sharing private raw data like their own experiences. By exchanging only learned models, MUs collectively compensate for individual limitations, and find more scalable, robust, and efficient task participation strategies. Comprehensive evaluations on both synthetic and real-world datasets show that FDRL-PPO consistently outperforms benchmark algorithms in terms of task completion ratio, fairness in task completion, energy consumption, and number of conflicting proposals.

联邦学习强化学习众包

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。