让多模态推理模型实时展示视觉依据,边思考边输出证据。
Real-Time Visual Attribution Streaming in Thinking Model

- 通过学习注意力特征,轻量估算语义区域的因果影响。
- 在五个基准上达到与耗时方法相当的可信度。
- 适合需要实时可解释性的AI应用开发人员。
我们提出一种摊销式框架,实现多模态思维模型中实时视觉归因流。当这些模型从截图生成代码或从图像解答数学问题时,其长推理过程应基于视觉证据。然而验证这种依赖性困难:忠实的因果方法需代价高昂的重复反向传播或扰动,而原始注意力图虽可即时获取,但缺乏因果有效性。为此,我们引入一种摊销方法,直接从注意力特征编码的丰富信号中学习估算语义区域的因果效应。在五个多样化基准和四种思维模型上,该方法实现与详尽因果方法相当的可信度,同时支持视觉归因流,使用户可在模型推理过程中实时观察支撑证据,而非事后。结果表明,通过轻量学习而非暴力计算,即可实现多模态思维模型的实时、可信归因。
原文摘要 · Abstract (English)
We present an amortized framework for real-time visual attribution streaming in multimodal thinking models. When these models generate code from a screenshot or solve math problems from images, their long reasoning traces should be grounded in visual evidence. However, verifying this reliance is challenging: faithful causal methods require costly repeated backward passes or perturbations, while raw attention maps offer instant access, they lack causal validity. To resolve this, we introduce an amortized approach that learns to estimate the causal effects of semantic regions directly from the rich signals encoded in attention features. Across five diverse benchmarks and four thinking models, our approach achieves faithfulness comparable to exhaustive causal methods while enabling visual attribution streaming, where users observe grounding evidence as the model reasons, not after. Our results demonstrate that real-time, faithful attribution in multimodal thinking models is achievable through lightweight learning, not brute-force computation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。