arXiv:2510.03722cs.LGstat.ML2025-10

提出一种可解释且性能优越的强化学习方法,通过自适应正则化提升决策质量。

Balancing Interpretability and Performance in Reinforcement Learning: An Adaptive Spectral Based Linear Approach

  • 基于谱滤波设计线性强化学习模型,融合正则化与偏差-方差权衡
  • 理论证明参数估计与泛化误差接近最优,实验在真实数据上表现优异
  • 适合需要可解释性的管理决策场景,如推荐系统、广告投放

强化学习广泛应用于序列决策,可解释性与性能对实际应用至关重要。现有方法多侧重性能,依赖事后解释实现可解释性。本文提出一种以可解释性为导向、同时提升性能的强化学习方法。具体而言,基于谱滤波扩展了基于岭回归的方法,阐明正则化对控制估计误差的作用,并设计了由偏差-方差权衡原则指导的自适应正则化参数选择策略。理论分析建立了参数估计和泛化误差的近似最优界。在模拟环境及快手、淘宝的真实数据集上的大量实验表明,该方法在决策质量上优于或匹配现有基线。通过可解释性分析,展示了学习策略如何做出决策,从而增强用户信任。结果表明,该方法有望弥合强化学习理论与实际决策之间的差距,在管理场景中实现可解释性、准确性与适应性的统一。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has been widely applied to sequential decision making, where interpretability and performance are both critical for practical adoption. Current approaches typically focus on performance and rely on post hoc explanations to account for interpretability. Different from these approaches, we focus on designing an interpretability-oriented yet performance-enhanced RL approach. Specifically, we propose a spectral based linear RL method that extends the ridge regression-based approach through a spectral filter function. The proposed method clarifies the role of regularization in controlling estimation error and further enables the design of an adaptive regularization parameter selection strategy guided by the bias-variance trade-off principle. Theoretical analysis establishes near-optimal bounds for both parameter estimation and generalization error. Extensive experiments on simulated environments and real-world datasets from Kuaishou and Taobao demonstrate that our method either outperforms or matches existing baselines in decision quality. We also conduct interpretability analyses to illustrate how the learned policies make decisions, thereby enhancing user trust. These results highlight the potential of our approach to bridge the gap between RL theory and practical decision making, providing interpretability, accuracy, and adaptability in management contexts.

强化学习可解释性谱方法自适应正则化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。