arXiv:2602.04355cs.CL2026-02

视觉信息在工作记忆中表现不如文本,模型更依赖最近记忆。

Can Vision Replace Text in Working Memory? Evidence from Spatial n-Back in Vision-Language Models

  • 用空间n-back任务对比视觉与文本输入下的模型表现
  • 文本任务准确率和敏感度显著高于视觉任务
  • 模型实际依赖近期刺激而非指令的延迟匹配,适合研究多模态记忆机制

工作记忆是智能行为的核心,提供动态的工作空间以维持和更新任务相关信息。近期研究使用n-back任务探测大语言模型中的类工作记忆行为,但尚不清楚当信息以视觉而非文本形式呈现时,相同任务是否引发类似计算过程。我们评估了Qwen2.5和Qwen2.5-VL在控制性空间n-back任务中的表现,该任务以匹配的文字渲染或图像渲染网格呈现。结果显示,无论模型类型,文本条件下的准确率和d'均显著高于视觉条件。通过逐次试验的对数概率证据分析发现,名义上的2/3-back任务常未能反映指令设定的延迟,反而与基于近期性的比较对齐。此外,网格大小会改变刺激流中的近期重复结构,从而影响干扰模式和错误分布。这些结果推动了对多模态工作记忆的计算敏感性解释。

原文摘要 · Abstract (English)

Working memory is a central component of intelligent behavior, providing a dynamic workspace for maintaining and updating task-relevant information. Recent work has used n-back tasks to probe working-memory-like behavior in large language models, but it is unclear whether the same probe elicits comparable computations when information is carried in a visual rather than textual code in vision-language models. We evaluate Qwen2.5 and Qwen2.5-VL on a controlled spatial n-back task presented as matched text-rendered or image-rendered grids. Across conditions, models show reliably higher accuracy and d' with text than with vision. To interpret these differences at the process level, we use trial-wise log-probability evidence and find that nominal 2/3-back often fails to reflect the instructed lag and instead aligns with a recency-locked comparison. We further show that grid size alters recent-repeat structure in the stimulus stream, thereby changing interference and error patterns. These results motivate computation-sensitive interpretations of multimodal working memory.

工作记忆多模态视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。