解决大模型强化学习中奖励信号难以定位责任归属的问题
From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models
- 提出六诊断框架,系统分析责任分配的失效原因
- 发现仅靠文本历史无法确定奖励归属方向
- 提供可复用的责任卡工具,支持因果验证与证据追溯
大语言模型的强化学习越来越多依赖稀疏结果奖励,但这类奖励无法说明是哪个词元、推理步骤、工具调用、记忆操作或智能体导致了结果。这一责任分配(CA)问题在推理型和代理型强化学习中尤为突出,后者因环境交互引入状态转移非闭合、部分可观测性、有限回放、异构动作、中间验证弱以及代理耦合等挑战而加剧。本文整合2024年1月至2026年7月31日间发表的69篇论文:56种核心责任分配方法与13种相关或边界使能技术,基于92份去重筛选记录。保留原始按方法分类的粒度,并新增六诊断框架,将假设破缺映射至识别障碍、估计器与评估控制。对42篇核心论文进行来源定位全文审计,两名算法研究员独立盲评252个诊断单元,一致率达88.5%;各诊断项科恩卡帕系数介于0.543至0.909之间,主要家族一致性为42/42(卡帕=1.000)。除分类体系外,证明恢复状态对比可识别协议特异性因果差异,指出纯文本历史可能使信用归属符号也无法确定,并提出可复用的CA-ID卡,关联每项主张与其估计量、证据来源与可证伪测试。原子级报告审计涵盖比较对象、预算对齐、消融实验、开销、不确定性及回放覆盖,不构建跨论文排行榜。配套仓库包含动态目录与决策辅助工具;冻结审计包将作为独立版本发布,区别于最小arXiv源文件。
原文摘要 · Abstract (English)
Reinforcement learning (RL) for large language models (LLMs) increasingly relies on sparse outcome rewards, yet such rewards say little about which token, reasoning step, tool call, memory operation, or agent caused an outcome. This credit assignment (CA) problem spans reasoning RL and becomes sharper in agentic RL, where environment interaction introduces transition non-closure, partial observability, limited replay, heterogeneous actions, weak intermediate verifiability, and agent coupling. We synthesize a unified corpus of 69 papers published from January 2024 through July 31, 2026: 56 core CA methods and 13 adjacent or boundary enablers, selected from 92 deduplicated screening records. We retain the original granularity-by-methodology taxonomy and add a six-diagnostic framework mapping assumption breaks to identification barriers, estimators, and evaluation controls. A source-located full-text audit covers a fixed 42-core-paper subset. Two algorithm researchers independently and blindly cross-coded 252 diagnostic cells, agreeing on 223 (88.5%); per-diagnostic Cohen's kappa ranges from .543 to .909, and principal-family agreement is 42/42 (kappa=1.000). Beyond taxonomy, we establish when restored-state comparisons identify a protocol-specific causal contrast, show that text-only histories can leave even the sign of credit unidentified, and introduce a reusable CA-ID Card linking each claim to its estimand, evidence provenance, and falsification test. An atomic reporting audit describes comparator, budget parity, ablation, overhead, uncertainty, and replay coverage without constructing a cross-paper leaderboard. The companion repository hosts a living catalog and decision aids; a dated release of the frozen audit bundle is planned there separately from the minimal arXiv source.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。