修复视频模型在多模态训练中丢失的时序推理能力
Lost in Adaptation: Layer-Selective Recovery of Temporal Reasoning in Video-Language Models

- 通过选择性合并不同层的模型参数来恢复时序推理
- 在5个数据集上提升时序推理能力,最高达27.8%相对增益
- 无需重新训练,适合已部署模型的后处理修复
多模态适应会削弱视频-语言模型(VLM)的时序推理(TR)能力,导致模型虽能感知关键事件,却无法推断其时间与因果结构。我们提出MERIT,一种无梯度的框架,通过分层选择性模型合并来修复该能力。MERIT为每个自注意力层分配VLM主导或LLM主导的插值,并使用协方差矩阵自适应进化策略(CMA-ES)在组合空间中搜索,以奖励时序推理提升并惩罚时序感知退化。在三个VLM家族和五个视频基准上,MERIT持续提升时序推理且保持时序感知;在紧凑诊断集上选出的方案可迁移至四个未见基准,相对增益最高达27.8%。干预性掩码与帧级归因表明,所选层对推理功能至关重要,且MERIT使决策更依赖于时空分布、因果相关的证据。这些结果确立了分层选择性合并作为修复多模态适配中时序推理退化的实用后处理机制,无需学习新参数。
原文摘要 · Abstract (English)
Multimodal adaptation can erode temporal reasoning (TR) in video-language models (VLMs), leaving models able to perceive salient events yet unable to infer their temporal and causal structure. We introduce MERIT, a gradient-free framework that repairs this capability through layer-selective model merging. MERIT assigns each self-attention layer a VLM-dominant or LLM-dominant interpolation and uses the Covariance Matrix Adaptation Evolution Strategy (CMA-ES) to search the resulting combinatorial space under an objective that rewards TR gains while penalizing temporal perception (TP) degradation. Across three VLM families and five video benchmarks, MERIT consistently improves TR while preserving TP; recipes selected on a compact diagnostic set transfer to four unseen benchmarks, with relative gains of up to 27.8%. Interventional masking and frame-level attribution further show that the selected layers are functionally important for reasoning and that MERIT shifts decisions toward temporally distributed, causally relevant evidence. These results establish layer-selective merging as a practical post-hoc mechanism for repairing video temporal reasoning degraded during multimodal adaptation, without learning new parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。