检验视觉记忆何时可安全遗忘,发现当前注意力低不等于未来不用。
Seen, Said, or Forgotten? A Causal Audit of Visual KV Memory Across Dialog Turns

- 设计因果审计框架,测试不同信息丢失对后续回答的影响。
- 当前注意力低的视觉区域,未来仍可能有用,随机选择反而更优。
- 已陈述的事实可通过文本记忆保留,未陈述的则无法可靠替代。
具有状态的多模态助手在初次编码图像后,可在多个对话轮次中回答相关问题。现有基于注意力的视觉键值(KV)淘汰机制假设:当前无关的信息未来也无用,但未来问题未知。本文提出因果视觉记忆审计(CVMA),通过配对单前填框架,测试当某一视觉区域、整张图像或先前助手文本不可用时,后续回答损失程度。在VisDial和ConvBench数据集上,当前注意力低的视觉区域,其未来价值排名甚至低于随机水平,尽管诊断性边际效用分析显示仍有较大选择空间。总体得分掩盖了这一失败,因为后期轮次未必依赖视觉;控制与生成的历史揭示第二种逃逸路径:已陈述的事实可由文本KV替代,但未陈述的事实无法可靠维持。实验表明,安全遗忘仅源于未来视觉依赖度低或事实被特定语言表达——而非当前注意力低。
原文摘要 · Abstract (English)
Stateful multimodal assistants encode an image once but may answer questions about it many turns later. Attention-guided visual-KV eviction assumes that evidence irrelevant now will remain dispensable, although future questions are unknown. We ask when a visual fact is actually safe to forget and introduce the Causal Visual Memory Audit (CVMA), a paired single-prefill framework that tests what later answers lose when a visual region, the whole image, or prior assistant text becomes unavailable. On VisDial and ConvBench, current attention can rank future-useful regions worse than random even though a diagnostic marginal-utility control shows substantial selection headroom. Aggregate scores hide this failure when later turns do not need vision; controlled and stock-generated histories reveal a second escape route, in which assistant-text KV replaces image KV for facts already stated but not reliably for unstated facts. In the tested stacks, safe forgetting is supported by low future visual dependence or fact-specific verbalization---not by low current attention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。