38B参数多模态模型,用轻量投影实现图文协同推理。
Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought
- 轻量视觉投影器实现无需重训的跨模态迁移。
- 混合优化策略提升图文对齐,推理链长度自适应优化。
- 在多个评测中表现优异,适合多模态推理研究者使用。
我们提出Skywork R1V,一个将R1系列大语言模型扩展至视觉模态的多模态推理模型,采用高效多模态迁移方法。通过轻量级视觉投影器,实现无需重训基础语言模型或视觉编码器的无缝多模态适配。为强化图文对齐,提出结合迭代监督微调(SFT)与组相对策略优化(GRPO)的混合优化策略,显著提升跨模态融合效率。此外,引入自适应长度的思维链蒸馏方法生成推理数据,动态优化推理链长度,提升推理效率并避免过度推理。实证评估显示,尽管仅含380亿参数,该模型在MMMU基准上取得69.0分,在MathVista上达67.5分;同时保持强文本推理能力,于AIME和MATH500分别取得72.0和94.0分。模型权重已公开,以促进开放与可复现性。
原文摘要 · Abstract (English)
We introduce Skywork R1V, a multimodal reasoning model extending the an R1-series Large language models (LLM) to visual modalities via an efficient multimodal transfer method. Leveraging a lightweight visual projector, Skywork R1V facilitates seamless multimodal adaptation without necessitating retraining of either the foundational language model or the vision encoder. To strengthen visual-text alignment, we propose a hybrid optimization strategy that combines Iterative Supervised Fine-Tuning (SFT) with Group Relative Policy Optimization (GRPO), significantly enhancing cross-modal integration efficiency. Additionally, we introduce an adaptive-length Chain-of-Thought distillation approach for reasoning data generation. This approach dynamically optimizes reasoning chain lengths, thereby enhancing inference efficiency and preventing excessive reasoning overthinking. Empirical evaluations demonstrate that Skywork R1V, with only 38B parameters, delivers competitive performance, achieving a score of 69.0 on the MMMU benchmark and 67.5 on MathVista. Meanwhile, it maintains robust textual reasoning performance, evidenced by impressive scores of 72.0 on AIME and 94.0 on MATH500. The Skywork R1V model weights have been publicly released to promote openness and reproducibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。