统一强化学习微调框架,支持多种训练模式和高效交互。
Trinity-RFT: A General-Purpose and Unified Framework for Reinforcement Fine-Tuning of Large Language Models
- 模块化设计,统一同步/异步、在线/离线等训练模式
- 高效鲁棒的智能体-环境交互接口,支持多样化应用
- 适合研究者快速构建强化学习系统,降低开发门槛
Trinity-RFT 是一个通用、统一且易用的大型语言模型强化微调(RFT)框架。其采用模块化解耦设计,包含三个核心组件:(1) RFT-core,统一并泛化了同步/异步、在线/离线、策略/非策略等多种RFT训练模式;(2) 高效鲁棒的智能体-环境交互集成;(3) 针对RFT优化的数据流水线。该框架可灵活适配多种应用场景,作为宏观与微观层面先进强化学习范式的统一研发平台。本文报告了其愿景、特性、设计与实现,并通过大量示例、应用与实验展示了其功能性和易用性。
原文摘要 · Abstract (English)
Trinity-RFT is a general-purpose, unified and easy-to-use framework designed for reinforcement fine-tuning (RFT) of large language models. It is built with a modular and decoupled design, consisting of (1) an RFT-core that unifies and generalizes synchronous/asynchronous, on-policy/off-policy, and online/offline modes of RFT; (2) seamless integration for agent-environment interaction with high efficiency and robustness; and (3) systematic data pipelines optimized for RFT. Trinity-RFT can be easily adapted for diverse application scenarios, and serves as a unified platform for development and research of advanced reinforcement learning paradigms at both macroscopic and microscopic levels. This technical report outlines the vision, features, design and implementations of Trinity-RFT, accompanied by extensive examples, applications and experiments that demonstrate its functionalities and user-friendliness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。