用强化学习+测试时扩增,少用数据也能高效推理视频。
Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning
- 纯强化学习训练,无需大量标注数据和微调
- 仅用3.6%样本提升2.4%准确率,视频-霍姆斯上提升4.2%
- 适合资源有限但需强视频推理能力的场景
尽管基于大语言模型(LLMs)的强化学习(RL)在视频推理方面取得进展,但数据收集与微调仍面临挑战。现有方法通常依赖大规模监督微调(SFT)及长链式思维(CoT)标注,成本高且难扩展。为此,我们提出Video-RTS,通过结合数据高效的强化学习与自适应视频测试时扩增(TTS)策略,显著提升数据效率。基于对数据缩放规律的观察,跳过耗资源的SFT步骤,采用仅需输出奖励的纯强化学习训练,无需额外标注或大规模微调。此外,引入稀疏到密集的视频TTS策略,通过迭代添加帧以提升输出一致性来优化推理。在多个视频推理基准上验证,Video-RTS仅使用3.6%训练样本即实现2.4%的准确率提升,在Video-Holmes上更达4.2%。纯强化学习与自适应视频TTS互补,共同支撑其强大推理性能。
原文摘要 · Abstract (English)
Despite advances in reinforcement learning (RL)-based video reasoning with large language models (LLMs), data collection and fine-tuning remain significant challenges. These methods often rely on large-scale supervised fine-tuning (SFT) with extensive video data and long Chain-of-Thought (CoT) annotations, making them costly and hard to scale. To address this, we present Video-RTS, a new approach to improve video reasoning capability with drastically improved data efficiency by combining data-efficient RL with a video-adaptive test-time scaling (TTS) strategy. Building on observations about the data scaling, we skip the resource-intensive SFT step and employ efficient pure-RL training with output-based rewards, requiring no additional annotations or extensive fine-tuning. Furthermore, to utilize computational resources more efficiently, we introduce a sparse-to-dense video TTS strategy that improves inference by iteratively adding frames based on output consistency. We validate our approach on multiple video reasoning benchmarks, showing that Video-RTS surpasses existing video reasoning models by 2.4% in accuracy using only 3.6% training samples. Specifically, Video-RTS achieves a 4.2% improvement on Video-Holmes, a recent and challenging video reasoning benchmark. Notably, our pure RL training and adaptive video TTS offer complementary strengths, enabling Video-RTS's strong reasoning performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。