挑战主流假设:无需大模型路由,也能高效持续学习多模态模型。
Is Our Benchmark Enough? An Analysis of Continual Learning for MLLMs

- 用冻结特征+任务原型实现无须训练的轻量级路由。
- 共享专家对持续学习无实质帮助,实测效果未提升。
- 现有基准存在任务太易分、顺序单一等缺陷,需重构评测标准。
持续适应对部署于动态领域的多模态大模型(MLLM)至关重要,但当前最先进的MR-LoRA方法高度依赖基于MLLM的路由器处理复杂多模态输入。本文在MLLM-CL基准上重新审视该假设,提出两点主张:首先,路由无需使用MLLM——一种无需训练、无需重放的原型路由方法(RePRo),仅通过冻结预训练特征与任务原型,即可在远低于MR-LoRA的计算成本下达到相当性能;其次,尽管理论上吸引人,共享专家并未提升MLLM的持续学习表现。这些发现源于MLLM-CL基准的两个结构性局限:(1)任务在表征空间中高度可分;(2)固定任务顺序导致结论对单一课程路径敏感,缺乏跨多样化持续学习轨迹的鲁棒性。因此,该基准更侧重孤立学习而非真正的持续迁移。这促使未来应设计新基准,包含重叠任务流形、多任务顺序、细粒度领域漂移及奖励前向迁移与保留能力的评估协议。
原文摘要 · Abstract (English)
Continual adaptation is essential for multimodal large language models (MLLMs) deployed across evolving domains, but the state-of-the-art MR-LoRA method highly relies on the assumption that a MLLM-based router is necessary to process complex multimodal inputs. This paper revisits this claim on the MLLM-CL benchmark and argues for two claims. \textbf{First}, routing does not require an MLLM: a simple training-free, replay-free ptotypical routing method (\textsc{RePRo}), uses frozen pretrained features and task prototypes to match the MLLM-based router of MR-LoRA at far lower computational cost. \textbf{Second}, shared experts do not improve continual learning for MLLMs, despite their theoretical appeal. We show that these findings arise from two structural limitations of MLLM-CL: (1) its tasks are \textbf{highly separable} in representation space, and (2) its fixed task order makes conclusions \textbf{sensitive to a single curriculum} rather than robust across diverse continual-learning trajectories. As a result, the benchmark primarily rewards learning in isolation rather than genuine continual transfer. This motivates a new design for future benchmarks of continual MLLM learning, with overlapping task manifolds, multiple task orders, fine-grained domain shifts, and evaluation protocols that reward forward transfer as well as retention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。