arXiv:2606.24622cs.AIcs.HC2026-06

让强化学习更安全可解释,用人类反馈训练出媲美真实奖励的模型。

Themis: An explainable AI-enabled framework for Reinforcement Learning with Human Feedback

论文配图:Themis: An explainable AI-enabled framework for Reinforcement Learning with Human Feedback
图 1 · 摘自论文原文
  • 融合可解释AI与人类反馈,构建可透明评估的RL训练框架。
  • 用人类偏好训练的奖励模型性能媲美甚至超过环境真实奖励信号。
  • 支持千人并发实验,云平台易用且自动扩展,适合大规模研究。

训练安全的强化学习系统本质上具有挑战性,无法保证避免不良行为。最有效的防御策略是(i)通过可解释性实现透明化,以及(ii)通过人类反馈实现对齐。尽管两者均表现良好,但目前尚无公开可用的框架能同时整合二者。为此,我们提出Themis,一个面向人类反馈强化学习的XAI增强型测试与评估框架。Themis支持超过200个常用环境,可轻松配置用于RL、透明性与对齐性实验。结果表明,Themis能够训练出匹配或超越环境真实奖励信号的奖励模型。我们还提供一个基于云的平台,用于收集人类反馈并管理实验,具备用户友好、自动扩展特性,可在无需额外开发成本的情况下支持多实验、大规模参与者。测试显示,单台普通商用机器即可支持千名用户连续进行实验。

原文摘要 · Abstract (English)

Training safe Reinforcement Learning (RL) systems is inherently challenging, with no guarantee of avoiding unwanted behaviors. The most effective defenses against this are (i) transparency through explainability and (ii) alignment via human feedback. While both show promising results, no publicly available framework currently combines them. To address this, we introduce Themis, an XAI-enabled testing and evaluation framework for Reinforcement Learning from Human Feedback. Themis supports over 200 widely used environments and is easily configurable for experiments in RL, transparency, and alignment. Our results show that Themis can train reward models that match or outperform the environment's true reward signal using human preferences. We also provide a cloud-based platform for collecting human feedback and managing experiments. It is user-friendly, auto-scalable, and supports large participant groups across multiple experiments without extra development overhead. Tests show Themis can support one thousand users in back-to-back experiments on a modest commercial machine.

强化学习可解释AI人类反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。