arXiv:2509.24713cs.LGcs.AI2025-09

通过识别神经回路提升奖励模型对罕见情况的鲁棒性

Circuit-Aware Reward Training: A Mechanistic Framework for Longtail Robustness in RLHF

  • 基于神经回路分析,定位奖励模型处理罕见事件的专用通路
  • 在长尾分布上显著降低奖励劫持风险,提升泛化能力
  • 适合关注对齐安全与鲁棒训练的研究者

基于人类反馈的强化学习(RLHF)奖励模型在长尾分布上存在系统性失效,导致奖励劫持和对齐偏差。我们提出一种机制可解释性框架,识别奖励模型中负责稀有事件处理的专用神经回路。受语言模型中稀有标记分布式特化的最新研究启发,我们假设奖励模型也发展出针对长尾场景的功能性独立回路。理论框架建立了回路专业化、奖励泛化界与长尾性能之间的形式联系。我们引入【Circuit-Aware Reward Training (CART)】,利用回路分析指导数据增强、正则化与集成策略。该方法既提供对奖励模型失败机制的理论洞察,也给出提升长尾鲁棒性的实用干预方案。

原文摘要 · Abstract (English)

Reinforcement Learning from Human Feedback (RLHF) reward models exhibit systematic failures on longtail distributions, leading to reward hacking and misalignment. We propose a mechanistic interpretability framework that identifies specialized neural circuits responsible for rare-event processing in reward models. Drawing from recent advances showing distributed specialization for rare tokens in language models\citep{liu2025no, liu2025emergent}, we hypothesize that reward models also develop functionally distinct circuits for longtail scenarios. Our theoretical framework establishes formal connections between circuit specialization, reward generalization bounds, and longtail performance. We introduce \textbf{Circuit-Aware Reward Training (CART)}, which uses circuit analysis to guide data augmentation, regularization, and ensemble strategies. This approach provides both theoretical insights into reward model failures and practical interventions for improving longtail robustness.

RLHF长尾鲁棒性可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。