让AI模型学会分辨哪些视觉修正来自真实证据,提升多模态学习准确性。
VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation

- 通过对比有无视觉证据时的预测变化,分离出真正由视觉支持的修正
- 在6个视觉任务上,40亿到90亿参数模型均优于传统方法
- 适合需要精准视觉推理的多模态应用,如图像描述生成
多模态在线策略蒸馏(OPD)通过特权视角教师监督学生生成轨迹来传递细粒度视觉知识。然而其下一个词修正常混合视觉信号、语言先验和教师特异性影响。核心挑战在于判断哪些修正由视觉证据支持,而非仅关注如何或何处蒸馏。本文提出视觉归因蒸馏(VAD),一种反事实目标重构算法,用于估计教师修正中可归因于视觉证据的部分。在每个学生生成前缀处,VAD对同一固定教师分别评估相关证据存在与缺失的情况。中心化对数概率的变化定义了ut,作为视觉证据方向的带符号代理,反映证据对候选词的支持或反驳程度。VAD将原修正投影至该代理,得到与干预对齐的成分和代理未解释残差,再从对齐成分重构以学生为中心的目标。训练中,此重构目标提供主要监督信号,而特权教师仅起弱正则作用。在6个细粒度视觉基准测试中,40亿和90亿参数规模下,VAD均超越直接特权视图蒸馏与视觉优势加权。逐词分析与受控目标实验表明,对齐成分富含任务相关的视觉修正,且在证据反驳错误答案时产生更强目标偏移。结果支持反事实目标重构作为源混合监督的有效替代方案。
原文摘要 · Abstract (English)
Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view teacher. Yet its next-token corrections are source-mixed, combining visual signals with linguistic priors and teacher-specific effects. The key challenge is to estimate which corrections are supported by visual evidence, not merely where or how strongly to distill. We introduce Visual Attribution Distillation (VAD), a counterfactual target-reconstruction algorithm that estimates the visually attributable part of a teacher correction. At each student-generated prefix, VAD evaluates the same fixed teacher with the relevant evidence present and removed. The corresponding change in centered log-probabilities defines ut, a signed proxy for the visual evidence direction that estimates how revealing the evidence supports or refutes candidate tokens. VAD projects the original correction onto this proxy to obtain an intervention-aligned component and a proxy-unexplained residual, then reconstructs a student-anchored target from the former. During training, this reconstructed target supplies the primary supervision signal, while the privileged teacher contributes a weak regularizer. Across six fine-grained visual benchmarks at 4B and 9B scales, VAD outperforms direct privileged-view distillation and visual-advantage weighting. Token- level and controlled-target analyses show that the proxy-aligned component is enriched in task-relevant visual corrections and yields stronger target shifts, especially when evidence refutes a mistaken answer. These results support counterfactual target reconstruction as an effective alternative to source-mixed supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。