arXiv:2603.04768cs.LG2026-03

用信息瓶颈与风险敏感强化学习,实现高速内存均衡器的高效鲁棒优化。

Distributional Reinforcement Learning with Information Bottleneck for Uncertainty-Aware DRAM Equalization

  • 结合信息瓶颈压缩信号,速度比传统眼图快51倍
  • 在4/8抽头配置下,最差情况性能提升33.8%/38.2%以上
  • 可量化不确定性,适合生产级内存系统部署

高速内存系统中均衡器参数优化对信号完整性至关重要。现有方法存在眼图评估计算开销大、优化目标为期望而非最坏情况、缺乏部署决策不确定性量化等问题。本文提出一种融合信息瓶颈隐表示与条件风险价值(CVaR)优化的分布式风险敏感强化学习框架。通过率失真最优信号压缩,在8个内存单元共240万波形上实现比眼图快51倍的加速,同时利用蒙特卡洛丢弃法量化认知不确定性。基于分位数回归的分布强化学习支持显式最坏情况优化,而PAC-Bayesian正则化提供泛化保证。实验表明,4抽头和8抽头配置下平均性能分别提升37.1%和41.5%,最坏情况分别达33.8%和38.2%,相较Q-learning基线提高80.7%和89.1%。框架实现62.5%高可靠性分类,基本消除多数配置的手动验证需求。结果表明该框架为生产规模均衡器优化提供了具备最坏情况保证的实用方案。

原文摘要 · Abstract (English)

Equalizer parameter optimization is critical for signal integrity in high-speed memory systems operating at multi-gigabit data rates. However, existing methods suffer from computationally expensive eye diagram evaluation, optimization of expected rather than worst-case performance, and absence of uncertainty quantification for deployment decisions. In this paper, we propose a distributional risk-sensitive reinforcement learning framework integrating Information Bottleneck latent representations with Conditional Value-at-Risk optimization. We introduce rate-distortion optimal signal compression achieving 51 times speedup over eye diagrams while quantifying epistemic uncertainty through Monte Carlo dropout. Distributional reinforcement learning with quantile regression enables explicit worst-case optimization, while PAC-Bayesian regularization certifies generalization bounds. Experimental validation on 2.4 million waveforms from eight memory units demonstrated mean improvements of 37.1\% and 41.5\% for 4-tap and 8-tap equalizer configurations with worst-case guarantees of 33.8\% and 38.2\%, representing 80.7\% and 89.1\% improvements over Q-learning baselines. The framework achieved 62.5\% high-reliability classification eliminating manual validation for most configurations. These results suggest the proposed framework provides a practical solution for production-scale equalizer optimization with certified worst-case guarantees.

强化学习内存优化不确定性量化分布式学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。