arXiv:2607.11312cs.CV2026-07

首个测试视频大模型从长视频中学习技能并实时应用的基准

SLVMBench: Skill Learning from Video Memory

论文配图:SLVMBench: Skill Learning from Video Memory
图 1 · 摘自论文原文
  • 用2-3小时长视频流测试模型记忆与提取操作知识的能力
  • 现有视频大模型在长视频中学习技能时表现显著下降
  • 适合研究实时技能获取与长时记忆的AI系统开发者

我们提出技能学习从视频记忆(SLVMBench),这是首个联合评估视频大语言模型能否从长视频记忆中学习技能并应用于实时任务的基准。该基准提供包含教程视频的2-3小时视频流,其中嵌入大量无关视频,模拟真实人类学习场景。要求模型基于所学技能回答正在进行视频中的实时问题。不同于强调被动理解的长视频理解基准,或依赖短时演示的技能学习基准,SLVMBench测试完整的知识记忆、提取与实时迁移流程。严格的人工标注包括亚秒级时间校准、人工设计的问题以排除常识猜测,并整合教程确保技能覆盖。对主流专有和开源视频大模型的评估显示,模型在从长视频中学习并应用技能方面表现不佳,且当技能信息置于长视频记忆中时性能大幅下降。这些结果揭示了当前视频大模型的关键局限,并确立了SLVMBench作为首个研究长上下文视频记忆中实时技能获取与应用的基准。

原文摘要 · Abstract (English)

We introduce Skill Learning from Video Memory (SLVMBench), the first benchmark that jointly evaluates whether video large language models (video-LLMs) can learn skills from long video memory and apply them to real-time tasks. SLVMBench presents models with 2-3 hour video streams that contain a tutorial video embedded in a stream of arbitrary irrelevant videos, resembling real-world human learning practices. Video-LLMs are asked to apply the acquired skill to answer real-time questions about an ongoing video. Unlike long-video understanding benchmarks that emphasize passive comprehension and skill-learning benchmarks that rely on short, immediate demonstrations, SLVMBench tests the full pipeline of memorizing and extracting procedural knowledge, as well as transferring it to real-time tasks. Moreover, rigorous human annotations feature sub-second-level temporal calibration, manually engineered questions eliminating common-sense guessing, and collated tutorials to ensure coverage of the required skills. Evaluations on state-of-the-art proprietary and open-source video LLMs show that video-LLMs struggle substantially with learning and applying skill knowledge from videos. Moreover, performance degrades markedly when the skill knowledge is placed within a long video memory. These results reveal a key limitation of existing video LLMs and position SLVMBench as the first benchmark for studying real-time skill acquisition and application from long-context video memory.

视频理解技能学习长视频大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。