arXiv:2506.21367cs.LGcs.AI2025-06被引 2

用图像增强正则化Q值分布,提升像素级强化学习性能

rQdia: Regularizing Q-Value Distributions With Image Augmentation

  • 通过MSE损失使增强图像下的Q值分布更均匀
  • 在MuJoCo上使DrQ、SAC分别在9/12、10/12任务中超越基线
  • 显著提升样本效率和长期训练表现,适合像素级连续控制研究

rQdia 在基于像素的深度强化学习中,利用图像增强正则化Q值分布。通过一个简单的辅助损失,以均方误差(MSE)使不同增强图像下的Q值分布趋于一致。该方法在MuJoCo连续控制套件的12个任务中,使DrQ和SAC分别在9/12和10/12任务上优于基线;在Atari Arcade环境的26个任务中,使Data-Efficient Rainbow在18/26任务上取得提升。性能增益体现在样本效率和长期训练稳定性上。此外,rQdia首次使无模型连续控制从像素输入超越状态编码基线。

原文摘要 · Abstract (English)

rQdia regularizes Q-value distributions with augmented images in pixel-based deep reinforcement learning. With a simple auxiliary loss, that equalizes these distributions via MSE, rQdia boosts DrQ and SAC on 9/12 and 10/12 tasks respectively in the MuJoCo Continuous Control Suite from pixels, and Data-Efficient Rainbow on 18/26 Atari Arcade environments. Gains are measured in both sample efficiency and longer-term training. Moreover, the addition of rQdia finally propels model-free continuous control from pixels over the state encoding baseline.

强化学习Q值正则化图像增强连续控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。