arXiv:2601.07224cs.AIcs.LG2026-01ACL被引 2

按数据冲突程度自动分配给SFT或RL,提升大模型训练效率

Consolidation or Adaptation? PRISM: Disentangling SFT and RL Data via Gradient Concentration

  • 基于梯度几何结构判断数据冲突度,决定分流到SFT或RL
  • 在WebShop和ALFWorld上性能超越顶尖方法,节省3.22倍计算成本
  • 适合追求高效训练的大模型对齐研究者

尽管结合监督微调(SFT)与强化学习(RL)已成为训练大语言模型智能体的标准范式,但两阶段间数据分配的有效机制仍缺乏深入探索。当前策略多依赖表面启发式规则,难以诊断模型内在学习需求。由于SFT旨在通过模仿实现模式固化,而RL则通过探索驱动结构适应,若数据与功能角色错配,将引发严重优化干扰。本文提出PRISM,一种基于认知架构理论的动力学感知框架,根据数据与模型现有知识的冲突程度进行数据仲裁。通过分析梯度的空间几何结构,PRISM识别出空间集中度高的数据为高冲突信号,需由RL进行结构重构;而梯度分布稀疏的数据则分配至SFT以实现高效固化。在WebShop和ALFWorld上的大量实验表明,PRISM实现了帕累托改进,在超越现有最优混合方法的同时,计算成本降低最多达3.22倍。研究结果表明,依据内部优化状态解耦数据对实现可扩展且稳健的智能体对齐至关重要。

原文摘要 · Abstract (English)

While Hybrid Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL) has become the standard paradigm for training LLM agents, effective mechanisms for data allocation between these stages remain largely underexplored. Current data arbitration strategies often rely on surface-level heuristics that fail to diagnose intrinsic learning needs. Since SFT targets pattern consolidation through imitation while RL drives structural adaptation via exploration, misaligning data with these functional roles causes severe optimization interference. We propose PRISM, a dynamics-aware framework grounded in Schema Theory that arbitrates data based on its degree of cognitive conflict with the model's existing knowledge. By analyzing the spatial geometric structure of gradients, PRISM identifies data triggering high spatial concentration as high-conflict signals that require RL for structural restructuring. In contrast, data yielding diffuse updates is routed to SFT for efficient consolidation. Extensive experiments on WebShop and ALFWorld demonstrate that PRISM achieves a Pareto improvement, outperforming state-of-the-art hybrid methods while reducing computational costs by up to 3.22$\times$. Our findings suggest that disentangling data based on internal optimization regimes is crucial for scalable and robust agent alignment.

大模型训练数据分配强化学习梯度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。