arXiv:2512.00694cs.CV2025-12被引 7

提出视频语言理解持续学习的新框架,明确区分稳定与可变部分。

Affordance-First Decomposition for Continual Learning in Video-Language Understanding

  • 用不变的效用令牌构建共享基础,仅在查询需要时动态调整。
  • 在领域增量视频问答中达51.6%准确率,遗忘率仅-1.8%。
  • 适合资源受限且需隐私保护的持续学习场景。

视频-语言理解的持续学习日益重要,因数据、领域和查询风格不断变化。现有方法常混淆应保持稳定的部分与需适应的部分,依赖固定路由或容量,或需重放旧视频。本文提出效用优先分解(AFD):将视频映射为缓慢变化的效用令牌,形成共享的时间对齐基础;同时设计轻量级、查询路由、冲突感知调度器,仅在必要时集中适应并扩展容量。基础通过弱对齐和教师一致性稳定,训练采用仅问题重放。AFD在多种协议上达到领先性能:领域增量视频问答平均准确率51.6%,遗忘率-1.8%;ViLCo R@[email protected]为29.6%(MQ)和20.7%(NLQ),[email protected](VQ)达18.4%;时间增量iVQA准确率39.5%,遗忘率-1.6%。整体上,AFD实现了交互中心基础与针对性适应的显式可解释分离。

原文摘要 · Abstract (English)

Continual learning for video--language understanding is increasingly important as models face non-stationary data, domains, and query styles, yet prevailing solutions blur what should stay stable versus what should adapt, rely on static routing/capacity, or require replaying past videos. We aim to explicitly specify where stability lives and where plasticity should be focused under realistic memory and privacy constraints. We introduce Affordance-First Decomposition (AFD): videos are mapped to slowly varying affordance tokens that form a shared, time-aligned substrate, while a lightweight, query-routed, conflict-aware scheduler concentrates adaptation and grows capacity only when needed. The substrate is stabilized via weak alignment and teacher consistency, and training uses question-only replay. AFD achieves state-of-the-art across protocols: 51.6% average accuracy with -1.8% forgetting on domain-incremental VideoQA, ViLCo R@[email protected] of 29.6% (MQ) and 20.7% (NLQ) with 18.4% [email protected] (VQ), and 39.5% accuracy with -1.6% forgetting on time-incremental iVQA. Overall, AFD offers an explicit, interpretable split between a stable interaction-centered substrate and targeted adaptation.

持续学习视频理解效用分解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。