arXiv:2603.27971cs.LG2026-03

自动挑选最优原型,让强化学习更可解释且不降效

Principal Prototype Analysis on Manifold for Interpretable Reinforcement Learning

  • 从数据中自动选原型,不再依赖人工设定
  • 在标准环境上性能媲美现有可解释模型
  • 适合追求模型透明度的RL研究者

近年来,强化学习(RL)广泛应用于实时游戏求解和基于人类偏好数据微调大语言模型,显著提升了与用户期望的一致性。然而,随着模型复杂度指数级增长,系统可解释性面临严峻挑战。尽管计算机视觉和自然语言处理领域已发展出多种可解释性方法以阐明局部与全局推理模式,但其在强化学习中的应用仍受限。直接扩展这些方法往往难以在强化学习场景中兼顾可解释性与性能。原型封装网络(PW-Nets)近期展现出潜力,可在不牺牲原始黑箱模型效率的前提下提升强化学习的可解释性。然而,这类方法通常需人工定义参考原型,常依赖专家领域知识。本文提出一种无需人工干预的方法,可从可用数据中自动选择最优原型。初步实验在标准Gym环境中表明,该方法性能与现有PW-Nets相当,同时保持与原始黑箱模型竞争的水平。

原文摘要 · Abstract (English)

Recent years have witnessed the widespread adoption of reinforcement learning (RL), from solving real-time games to fine-tuning large language models using human preference data significantly improving alignment with user expectations. However, as model complexity grows exponentially, the interpretability of these systems becomes increasingly challenging. While numerous explainability methods have been developed for computer vision and natural language processing to elucidate both local and global reasoning patterns, their application to RL remains limited. Direct extensions of these methods often struggle to maintain the delicate balance between interpretability and performance within RL settings. Prototype-Wrapper Networks (PW-Nets) have recently shown promise in bridging this gap by enhancing explainability in RL domains without sacrificing the efficiency of the original black-box models. However, these methods typically require manually defined reference prototypes, which often necessitate expert domain knowledge. In this work, we propose a method that removes this dependency by automatically selecting optimal prototypes from the available data. Preliminary experiments on standard Gym environments demonstrate that our approach matches the performance of existing PW-Nets, while remaining competitive with the original black-box models.

强化学习可解释性原型学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。