arXiv:2607.13546cs.LG2026-07

考虑工人能力随经验提升的实时招募,用有限预算最大化感知效果。

From Novice to Expert: Cost-Aware Bandits for Evolving Worker Performance in Crowdsensing

论文配图:From Novice to Expert: Cost-Aware Bandits for Evolving Worker Performance in Crowdsensing
图 1 · 摘自论文原文
  • 设计自适应算法,同时学习工人能力增长与成本变化趋势。
  • 在预算耗尽前实现感知质量持续提升,比基线提升显著。
  • 适合动态人力调度场景,尤其对长期任务平台有实用价值。

移动众包感知(MC)通过智能手机招募用户完成感知任务,支持交通监控、环境感知等大规模应用。核心挑战在于预算受限下的在线工人招募,需在不确定性中学习工人感知性能。现有方法通常假设工人性能恒定,但现实中其能力会随经验提升并趋于稳定,且感知成本因设备和上下文状态变化而未知。本文研究一种预算约束的在线招募问题:每轮选择一名工人,观测其感知质量与实际成本,其中工人期望质量随参与次数增加而上升,最终趋于平台。我们将其建模为结构化多臂赌博机,每个工人的期望奖励是未知的先增后稳函数,成本亦未知。提出一个成本感知的在线学习框架,联合学习演化奖励轨迹与异质成本,检测性能饱和点,并优化预算分配以最大化长期感知效用。理论分析证明了性能保障,并通过大量实验验证该方法优于忽略经验动态或假设成本已知的基线。

原文摘要 · Abstract (English)

Mobile crowdsensing (MC) recruits mobile users to perform sensing tasks using their smartphones, enabling large-scale applications such as traffic monitoring and environmental sensing. A fundamental challenge is online worker recruitment under uncertainty, where the platform must learn workers' sensing performance while operating with a limited budget. Existing learning-based MC recruitment methods typically assume that each worker's sensing quality is stationary with a fixed mean over time. In practice, however, worker performance often improves with experience and eventually stabilizes, while the incurred sensing cost can be unknown in advance due to time-varying device and context states. In this paper, we study a budget-constrained online recruitment problem in which the platform selects one worker in each round, observes the sensing quality and incurred cost, where the expected sensing quality of each worker increases with experience and eventually converges to a plateau, and repeats until the budget is exhausted. We formulate this problem as a structured bandit model where each worker's expected reward evolves according to an unknown increasing-then-converging function of its participation count, and each worker has an unknown expected cost. We develop a cost-aware online learning framework that jointly learns evolving reward trajectories and heterogeneous costs, detects performance saturation, and allocates the limited budget to maximize long-term sensing utility. We provide theoretical performance guarantees and validate the proposed approach through extensive experiments, demonstrating consistent improvements over baselines that ignore experience-driven dynamics or assume known costs.

众包感知在线学习预算优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。