arXiv:2602.02408cs.CVcs.AI2026-02

用人类推理编辑视觉语言模型,提升纠错泛化能力

ReasonEdit: Editing Vision-Language Models using Human Reasoning

  • 让用户在编辑时提供推理过程,动态存储并检索相关事实
  • 在4个VLM上多个推理型图像问答数据集上达到顶尖编辑效果
  • 适合需要高精度、可解释性编辑的视觉语言模型应用

模型编辑旨在修正大预训练模型中的错误,同时不改变无关行为。尽管已有研究针对视觉语言模型(VLMs)进行编辑,但尚未有方法处理依赖推理的任务——这类任务通常需要人类和模型对图像进行推理分析。为此,我们提出ReasonEdit,首个支持用户在编辑过程中阐述推理逻辑的VLM编辑器,引入一种新的实用编辑范式。ReasonEdit持续将人类推理存入代码本,并在推理时通过受网络科学启发的拓扑平衡多模态嵌入方法,仅检索相关事实。在四个VLM及多个基于推理的图像问答数据集上,ReasonEdit均实现当前最优编辑性能,表明在编辑过程中引入人类推理能显著提升编辑泛化能力。

原文摘要 · Abstract (English)

Model editing aims to correct errors in large, pretrained models without altering unrelated behaviors. While some recent works have edited vision-language models (VLMs), no existing editors tackle reasoning-heavy tasks, which typically require humans and models to reason about images. We therefore propose ReasonEdit, the first VLM editor to let users explain their reasoning during editing, introducing a new, practical model editing setup. ReasonEdit continuously stores human reasoning in a codebook, and retrieves only relevant facts during inference using a novel topology-balanced multimodal embedding method inspired by network science. Across four VLMs on multiple rationale-based visual question answering datasets, ReasonEdit achieves state-of-the-art editing performance, ultimately showing that using human reasoning during editing greatly improves edit generalization.

模型编辑视觉语言人类推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。