arXiv:2602.08439cs.CV2026-02被引 8

让AI通过视频示范学习新任务,突破静态知识评估局限。

Demo-ICL: In-Context Learning for Procedural Video Knowledge Acquisition

  • 用视频或字幕示范引导模型动态学习新知识
  • 构建1200个教学视频的基准测试,挑战模型适应能力
  • 适合研究多模态模型上下文学习与视频理解的学者

尽管近期多模态大模型在视频理解方面能力提升显著,现有视频评测仍主要基于模型内部的静态知识,而非其从少量示例中动态学习与适应的能力。为此,我们提出「演示驱动的视频上下文学习」(Demo-ICL)任务,聚焦于通过上下文示范来理解目标视频并回答问题。同时,我们构建了 Demo-ICL-Bench 基准,涵盖1200段来自YouTube的教学视频及对应问题,从中提取两类示范:(i) 基于视频字幕的文本示范;(ii) 对应的教学视频作为视频示范。为有效应对该挑战,我们设计了双阶段训练策略的 Demo-ICL 模型:视频监督微调 + 信息辅助直接偏好优化,共同提升模型从上下文示例中学习的能力。大量实验表明,该基准具有挑战性,且 Demo-ICL 显著优于现有主流多模态大模型,揭示了未来研究方向。

原文摘要 · Abstract (English)

Despite the growing video understanding capabilities of recent Multimodal Large Language Models (MLLMs), existing video benchmarks primarily assess understanding based on models' static, internal knowledge, rather than their ability to learn and adapt from dynamic, novel contexts from few examples. To bridge this gap, we present Demo-driven Video In-Context Learning, a novel task focused on learning from in-context demonstrations to answer questions about the target videos. Alongside this, we propose Demo-ICL-Bench, a challenging benchmark designed to evaluate demo-driven video in-context learning capabilities. Demo-ICL-Bench is constructed from 1200 instructional YouTube videos with associated questions, from which two types of demonstrations are derived: (i) summarizing video subtitles for text demonstration; and (ii) corresponding instructional videos as video demonstrations. To effectively tackle this new challenge, we develop Demo-ICL, an MLLM with a two-stage training strategy: video-supervised fine-tuning and information-assisted direct preference optimization, jointly enhancing the model's ability to learn from in-context examples. Extensive experiments with state-of-the-art MLLMs confirm the difficulty of Demo-ICL-Bench, demonstrate the effectiveness of Demo-ICL, and thereby unveil future research directions.

视频理解上下文学习多模态演示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。