arXiv:2604.00677cs.CV2026-04被引 1

为视频大模型设计新基准,揭示持续学习的三大瓶颈。

CL-VISTA: Benchmarking Continual Learning in Video Large Language Models

  • 构建8个跨感知、理解、推理的任务,模拟真实数据分布变化。
  • 10种主流方法均在性能、效率、内存上存在根本权衡。
  • 适合研究多模态持续学习的算法设计与评估者参考。

视频大语言模型(Video-LLMs)需具备持续学习能力以适应非平稳现实数据。然而现有基准存在明显不足:多数仍基于无大规模预训练的模型,且普遍将单一数据集划分为子任务,导致任务冗余高,对预训练的Video-LLMs几乎不产生遗忘。为此,我们提出CL-VISTA,一个专为视频大模型持续学习设计的基准。通过整合8个涵盖感知、理解与推理的多样化任务,引发显著分布偏移,有效暴露灾难性遗忘。我们建立包含6种评估协议的综合框架,覆盖性能、计算效率与内存占用三个关键维度。其中性能维度引入通用视频理解评估,检验持续学习方法是否真正提升基础智能,而非仅实现任务特异性过拟合。对10种主流持续学习方法的广泛测试表明:无一方法在所有维度上表现最优。能有效缓解遗忘的方法往往牺牲泛化能力,或带来高昂的计算与内存开销。我们期望CL-VISTA为多模态基础模型的持续学习研究提供关键洞见。

原文摘要 · Abstract (English)

Video Large Language Models (Video-LLMs) require continual learning to adapt to non-stationary real-world data. However, existing benchmarks fall short of evaluating modern foundation models: many still rely on models without large-scale pre-training, and prevailing benchmarks typically partition a single dataset into sub-tasks, resulting in high task redundancy and negligible forgetting on pre-trained Video-LLMs. To address these limitations, we propose CL-VISTA, a benchmark tailored for continual video understanding of Video-LLMs. By curating 8 diverse tasks spanning perception, understanding, and reasoning, CL-VISTA induces substantial distribution shifts that effectively expose catastrophic forgetting. To systematically assess CL methods, we establish a comprehensive evaluation framework comprising 6 distinct protocols across 3 critical dimensions: performance, computational efficiency, and memory footprint. Notably, the performance dimension incorporates a general video understanding assessment to assess whether CL methods genuinely enhance foundational intelligence or merely induce task-specific overfitting. Extensive benchmarking of 10 mainstream CL methods reveals a fundamental trade-off: no single approach achieves universal superiority across all dimensions. Methods that successfully mitigate catastrophic forgetting tend to compromise generalization or incur prohibitive computational and memory overheads. We hope CL-VISTA provides critical insights for advancing continual learning in multimodal foundation models.

持续学习视频理解大模型评估多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。