arXiv:2606.05259cs.CV2026-06被引 1

构建首个知识与推理密集型视频理解数据集,提升模型深层理解能力。

VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding

论文配图:VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding
图 1 · 摘自论文原文
  • 基于专家领域视频设计人机协作生成流程,聚焦深度推理能力。
  • 包含31.5万条推理样本,显著提升知识密集型视频理解性能。
  • 适合关注视频推理、多模态理解与数据设计的研究者使用。

我们提出VideoKR,首个专为强化知识与推理密集型视频理解而设计的大规模训练语料库。它包含145,000个新收集的CC许可专家领域视频上的315,000个视频推理样例。我们开发了一种人机协同、技能导向的样例生成流程,旨在逐步提升视频推理深度,同时保证样例及其思维链(CoT)推理过程的难度、多样性和可靠性。我们还构建了VideoKR-Eval,一个由专家标注的新基准,其问题要求真正的视频理解与知识密集型推理,而非依赖文本捷径。实验表明,在标准SFT→GRPO流程下,经过VideoKR后训练的模型在知识密集型视频推理任务上优于先前方法,同时在通用视频推理任务上仍具竞争力,凸显数据设计对视频推理进展的关键作用。我们进一步进行了全面消融实验,为未来研究提供可操作的洞见。

原文摘要 · Abstract (English)

We introduce VideoKR, the first large-scale training corpus specifically designed to strengthen knowledge- and reasoning-intensive video understanding. It comprises 315K video reasoning examples over 145K newly collected, CC-licensed, expert-domain videos. We develop a human-in-the-loop, skill-oriented example generation pipeline that targets progressively deeper video reasoning capabilities while ensuring the difficulty, diversity, and reliability of both the examples and their CoT rationales. We also curate VideoKR-Eval, a new expert-annotated benchmark where questions require genuine video understanding and knowledge-intensive reasoning rather than textual shortcuts. Our experiments show that, under a standard SFT$\rightarrow$GRPO pipeline, models post-trained on VideoKR outperform prior post-training approaches on knowledge-intensive video reasoning while remaining competitive on general video reasoning, highlighting data design as a key driver of progress in video reasoning. We further conduct comprehensive ablations to isolate the contributions of VideoKR, providing actionable insights for future work.

视频理解知识推理数据集多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。