arXiv:2604.03562cs.AI2026-04

动态奖励反而降低卫星调度效率,稳定权重更优。

When Adaptive Rewards Hurt: Causal Probing and the Switching-Stability Dilemma in LLM-Guided LEO Satellite Scheduling

  • 用单变量因果探测法分析奖励项影响,发现切换惩罚增20%可提升157Mbps
  • 固定权重下速率342.1 Mbps,动态权重仅103.3±96.8 Mbps,性能反降
  • 适合需理解自然语言意图的场景,简单任务用传统方法即可

在多波束低轨卫星调度中,自适应奖励设计本应优于静态权重,但实验揭示了切换-稳定性悖论:近似恒定的奖励权重(342.1 Mbps)显著优于精心调校的动态权重(103.3±96.8 Mbps),因PPO需准平稳奖励信号以实现值函数收敛。权重自适应——无论质量如何——都会因反复重启收敛而损害性能。我们引入单变量因果探测法,独立扰动各奖励项±20%,在5万步后测量PPO响应。结果发现反直觉的杠杆效应:切换惩罚增加20%可使极地切换与热冷区段分别提升157、130 Mbps。在四种MDP架构(固定、规则、学习型MLP、微调LLM)上评估,已知与新流量场景下,MLP分别达357.9与325.2 Mbps,而微调LLM因权重振荡崩溃至45.3±43.0 Mbps,非知识不足,而是输出一致性才是关键约束。研究为通信系统中LLM-DRL融合提供实证路线图,明确其不可替代价值(自然语言意图理解)与替代场景。

原文摘要 · Abstract (English)

Adaptive reward design for deep reinforcement learning (DRL) in multi-beam LEO satellite scheduling is motivated by the intuition that regime-aware reward weights should outperform static ones. We systematically test this intuition and uncover a switching-stability dilemma: near-constant reward weights (342.1 Mbps) outperform carefully-tuned dynamic weights (103.3+/-96.8 Mbps) because PPO requires a quasistationary reward signal for value function convergence. Weight adaptation-regardless of quality-degrades performance by repeatedly restarting convergence. To understand why specific weights matter, we introduce a single-variable causal probing method that independently perturbs each reward term by +/-20% and measures PPO response after 50k steps. Probing reveals counterintuitive leverage: a +20% increase in the switching penalty yields +157 Mbps for polar handover and +130 Mbps for hot-cold regimes-findings inaccessible to human experts or trained MLPs without systematic probing. We evaluate four MDP architect variants (fixed, rule-based, learned MLP, finetuned LLM) across known and novel traffic regimes. The MLP achieves 357.9 Mbps on known regimes and 325.2 Mbps on novel regimes, while the fine-tuned LLM collapses to 45.3+/-43.0 Mbps due to weight oscillation rather than lack of domain knowledge-output consistency, not knowledge, is the binding constraint. Our findings provide an empirically-grounded roadmap for LLM-DRL integration in communication systems, identifying where LLMs add irreplaceable value (natural language intent understanding) versus where simpler methods suffice.

强化学习卫星调度因果探测奖励设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。