arXiv:2606.28719cs.AI2026-06

用双记忆系统让视觉语言模型动态适应新环境。

ComMem: Complementary Memory Systems for Test-Time Adaptation of Vision-Language Models

论文配图:ComMem: Complementary Memory Systems for Test-Time Adaptation of Vision-Language Models
图 1 · 摘自论文原文
  • 模仿海马体与皮层分工,分设快速记忆和慢速抽象记忆。
  • 在15个数据集上显著优于现有方法,跨域泛化能力更强。
  • 适合需要实时适应的视觉语言应用,如智能客服、机器人感知。

视觉语言模型(VLMs)在动态现实环境中的测试时适应(TTA)至关重要。然而,现有方法通常仅局部调整且无法积累长期知识,或仅依赖单一模态,未能利用VLM固有的多模态特性。受生物大脑互补记忆系统的启发,我们提出ComMem,模拟海马体与皮层的协同作用,实现高效的VLM TTA。ComMem包含两个核心组件:类海马体的快速适应详细记忆,从高置信度测试样本构建动态视觉缓存;类皮层的缓慢整合抽象记忆,持续优化全局文本原型。对于每个测试实例,ComMem联合优化两套记忆系统以确保跨模态一致性。在15个基准数据集上的大量实验表明,ComMem在自然分布偏移和跨数据集泛化场景下均显著优于当前最优方法,为提升VLM实际适应能力提供了新方向。

原文摘要 · Abstract (English)

Test-time adaptation (TTA) of vision-language models (VLMs) is essential for their robust deployment in dynamic, real-world environments. However, existing TTA methods often adapt locally without accumulating knowledge over time, or operating within a single modality without exploiting VLMs' inherently multi-modal nature. Inspired by the \textbf{Com}plementary \textbf{Mem}ory systems of the biological brain, we propose \textbf{ComMem}, an innovative approach that mimics the distinct but cooperative roles of the hippocampus and neocortex to enable effective TTA for VLMs. ComMem consists of two key components: a fast-adapting detailed memory, akin to the hippocampus, that forms a dynamic visual cache from high-confidence test samples; and a slow-integrating abstract memory, akin to the neocortex, that continually refines global textual prototypes. For each test instance, ComMem jointly optimizes both memory systems to ensure cross-modal consistency. Extensive experiments on 15 benchmark datasets show that ComMem significantly outperforms state-of-the-art methods under both natural distribution shifts and cross-dataset generalization, offering a promising direction for enhancing VLMs' practical adaptability.

视觉语言模型测试时适应双记忆系统跨域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。