为大模型强化微调设计自动故障管理框架,提升训练稳定性。
Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning

- 构建首个细粒度强化微调故障基准RFT-FaultBench,涵盖5类16种故障
- 发现故障在训练动态中可被观测且具独特特征指纹,验证可识别性
- 提出闭环自动故障管理框架RFT-FM,实现异常检测、诊断与修复一体化
强化微调(RFT)已成为大语言模型后训练的核心范式,但其训练过程仍极脆弱。现有工作多从系统层面提升可靠性或针对特定子问题修改算法,却忽视了训练过程级的故障管理。当前出现问题仍依赖人工排查,自动故障管理几乎空白。本文首次系统探索RFT的故障管理:构建首个细粒度故障基准RFT-FaultBench,覆盖5类故障家族、16种故障类型,包含779次训练运行、22,549条训练步记录和1,457,288条轨迹级数据。实证研究发现,RFT故障在训练动态中可观测,且具有可区分的故障指纹。基于此,提出RFT-FM框架,实现异常检测、故障诊断与自动修复的闭环管理。实验表明,该基准既非简单也未饱和,尤其在细微故障设置下仍具挑战;而RFT-FM在检测、诊断与缓解故障方面表现强劲。
原文摘要 · Abstract (English)
Reinforcement fine-tuning (RFT) has become a core paradigm for post-training large language models, yet its training process remains highly fragile. Existing efforts mainly improve reliability at the system level or address specific issues in individual subproblems by modifying RFT algorithms. Despite their effectiveness, they largely overlook the problem of failure management at the training-process level. When training goes wrong, practitioners still rely heavily on expert-driven manual inspection and correction, and automatic failure management for RFT remains largely unexplored. In this paper, we take a first step toward systematic failure management for reinforcement fine-tuning. To understand the empirical structure of RFT failures, we first construct RFT-FaultBench, the first benchmark for fine-grained failures in reinforcement fine-tuning, covering 5 fault families, 16 fault types, 779 training runs, 22,549 train-step records, and 1,457,288 trajectory-level records. Based on this benchmark, we conduct a comprehensive empirical study showing that RFT failures are both observable from training dynamics and distinguishable through their empirical fault fingerprints. Building on these findings, we propose RFT-FM, an automatic failure management framework for reinforcement fine-tuning that unifies anomaly detection, failure diagnosis, and auto remediation in a closed loop. Experimental results show that RFT-FaultBench is neither trivial nor saturated: it exhibits clear anomaly structure while still posing substantial challenges, especially under subtle fault settings. Moreover, RFT-FM shows strong capability in detecting, diagnosing, and mitigating RFT failures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。