针对社交平台定制大模型,提升小规模模型的适应性与稳定性。
RedOne 2.0: Rethinking Domain-specific LLM Post-Training in Social Networking Services
- 分三阶段训练:先探索、再针对性微调、最后强化学习优化
- 40亿参数模型平均性能比70亿基线高2.41点,用更少数据达8.74提升
- 适合资源有限但需快速适配社交语境的场景,如客服或内容生成
作为人际互动与信息传播的重要媒介,社交网络服务(SNS)对大语言模型(LLMs)提出独特挑战:负载异质性强、流行语和文化规范变化快、多语言多元语料导致显著分布偏移。监督微调(SFT)虽能提升专属性能,但常引发分布内表现与分布外鲁棒性之间的“跷跷板”效应,尤其在小型模型上更为明显。为此,我们提出RedOne 2.0,一种面向SNS的大型语言模型,采用渐进式、基于强化学习优先的后训练范式,实现快速且稳定的适应。该流程分为三个阶段:(1) 在精选SNS语料上进行探索性学习,建立初始对齐并识别系统性弱点;(2) 针对性微调,仅对诊断出的薄弱环节应用SFT,同时混合少量通用数据以缓解遗忘;(3) 精炼学习,重新引入以SNS为中心的强化学习信号,巩固改进并协调各任务间的权衡。在涵盖三类任务的多个评估中,我们的40亿参数模型相比70亿规模的次优基线平均提升2.41。此外,与以SFT为核心的RedOne方法相比,RedOne 2.0仅使用不到一半的数据即实现8.74的平均性能提升,证明其在紧凑规模下具备卓越的数据效率与稳定性。总体而言,RedOne 2.0为SNS场景下的领域专用大模型建立了具有竞争力且成本可控的基准,提升了能力而不牺牲鲁棒性。
原文摘要 · Abstract (English)
As a key medium for human interaction and information exchange, social networking services (SNS) pose unique challenges for large language models (LLMs): heterogeneous workloads, fast-shifting norms and slang, and multilingual, culturally diverse corpora that induce sharp distribution shift. Supervised fine-tuning (SFT) can specialize models but often triggers a ``seesaw'' between in-distribution gains and out-of-distribution robustness, especially for smaller models. To address these challenges, we introduce RedOne 2.0, an SNS-oriented LLM trained with a progressive, RL-prioritized post-training paradigm designed for rapid and stable adaptation. The pipeline consist in three stages: (1) Exploratory Learning on curated SNS corpora to establish initial alignment and identify systematic weaknesses; (2) Targeted Fine-Tuning that selectively applies SFT to the diagnosed gaps while mixing a small fraction of general data to mitigate forgetting; and (3) Refinement Learning that re-applies RL with SNS-centric signals to consolidate improvements and harmonize trade-offs across tasks. Across various tasks spanning three categories, our 4B scale model delivers an average improvements about 2.41 over the 7B sub-optimal baseline. Additionally, RedOne 2.0 achieves average performance lift about 8.74 from the base model with less than half the data required by SFT-centric method RedOne, evidencing superior data efficiency and stability at compact scales. Overall, RedOne 2.0 establishes a competitive, cost-effective baseline for domain-specific LLMs in SNS scenario, advancing capability without sacrificing robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。