利用视频中的关键帧信息,提升大模型推理能力。
Video-OPSD: Exploiting Privileged Visual Evidence for On-Policy Self-Distillation in Video Large Language Models

- 用标注的关键帧构建教师模型,聚焦重要视觉信息。
- 在多个基准上优于标准自蒸馏方法,性能接近GRPO。
- 适合需要高效训练的视频理解任务研究者。
近期,基于策略的自蒸馏(OPSD)通过来自特权自教师的密集标记级监督,有效提升了策略优化。尽管前景广阔,但该方法在视频大语言模型(Video-LLMs)中仍研究不足。现有方法通常通过增加上下文信息构建教师,而保持学生输入不变。然而,视频推理本身提供了独特的特权监督来源:长视频包含大量时间冗余,仅少数帧即可支持回答问题。本文提出 Video-OPSD,利用主输入内的关键视觉证据进行教师构建与知识迁移。首先,证据锚定自教师仅依赖标注证据帧,学生则仍基于完整视频推理,使教师提供更精准指导。其次,证据引导的标记优化根据每个推理标记对关键帧的依赖程度动态加权,强化感知基础推理。在多个视频理解与推理基准上的实验表明,Video-OPSD 在多种骨干网络上均优于标准 OPSD,性能接近 GRPO,且训练时间显著更短,为 Video-LLMs 提供了一种高效有效的后训练方案。
原文摘要 · Abstract (English)
On-policy self-distillation (OPSD) has recently emerged as an effective post-training paradigm that improves policy optimization through dense token-level supervision from a privileged self-teacher. Despite its promise, OPSD remains largely underexplored for Video Large Language Models (Video-LLMs). Existing methods typically construct privileged teachers by augmenting their context with additional information while keeping the primary input unchanged for both teacher and student. Video reasoning, however, offers a distinct source of privileged supervision within the primary input itself: long videos contain substantial temporal redundancy, and only a small subset of frames provides the evidence necessary to answer a question. Building on this observation, we present $\textbf{Video-OPSD}$, an OPSD framework that exploits privileged visual evidence for both self-teacher construction and knowledge transfer. First, our Evidence-Grounded Self-Teacher conditions the teacher exclusively on annotated evidence frames while the student continues to reason over the complete video. This focused visual input enables the teacher to provide more informative supervision. Second, our Evidence-Guided Token Optimization adaptively weights token-level distillation according to each reasoning token's reliance on privileged visual evidence, thereby emphasizing perceptually grounded reasoning. Experiments across video understanding and reasoning benchmarks show that $\textbf{Video-OPSD}$ consistently improves upon Standard OPSD across multiple backbones and achieves performance comparable to GRPO while requiring substantially less training time, establishing an effective and efficient post-training approach for Video-LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。