用强化学习让40亿参数小模型学会视频推理,还能解释思考过程。
TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning
- 基于40亿参数的小模型,通过强化学习提升视频理解与推理能力。
- 在通用视频问答数据集上显著提升表现,出现‘顿悟’式推理现象。
- 适合资源有限的研究者探索小模型的思维能力,提供实用经验。
近期,通过强化学习提升大型多模态模型(LMMs)的推理能力取得显著进展。然而,现有工作大多依赖数学与代码等高推理强度数据集,且通常以大规模模型为基础。我们认为,对计算资源有限的研究者而言,探索小规模模型的推理能力仍具价值。同时,使模型在通用问答数据集上具备解释其推理过程的能力同样重要。为此,我们提出小型视频推理模型TinyLLaVA-Video-R1。该模型基于参数不超过40亿的可追溯训练视频理解模型TinyLLaVA-Video,经过在通用视频问答数据集上的强化学习后,不仅显著提升了推理与思考能力,还展现出“顿悟时刻”的涌现特性。此外,我们分享了一系列实验发现,旨在为未来小规模模型中视频推理(思考)能力的探索提供实践启示。项目地址:https://github.com/ZhangXJ199/TinyLLaVA-Video-R1。
原文摘要 · Abstract (English)
Recently, improving the reasoning ability of large multimodal models (LMMs) through reinforcement learning has made great progress. However, most existing works are based on highly reasoning-intensive datasets such as mathematics and code, and researchers generally choose large-scale models as the foundation. We argue that exploring small-scale models' reasoning capabilities remains valuable for researchers with limited computational resources. Moreover, enabling models to explain their reasoning processes on general question-answering datasets is equally meaningful. Therefore, we present the small-scale video reasoning model TinyLLaVA-Video-R1. Based on TinyLLaVA-Video, a traceably trained video understanding model with no more than 4B parameters, it not only demonstrates significantly improved reasoning and thinking capabilities after using reinforcement learning on general Video-QA datasets, but also exhibits the emergent characteristic of "aha moments". Furthermore, we share a series of experimental findings, aiming to provide practical insights for future exploration of video reasoning (thinking) abilities in small-scale models. It is available at https://github.com/ZhangXJ199/TinyLLaVA-Video-R1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。