arXiv:2608.26993cs.CV2026-08

提出Aphanta框架,评估图像编辑在多模态推理中的实际效用

Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning

论文配图:Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning
图 1 · 摘自论文原文
  • 构建闭环诊断框架,对比三种推理条件下的表现
  • 发现编辑效果受任务类型影响显著,视觉线索注入最有效
  • 为多模态模型与编辑器协作提供可复用的评估协议

显式视觉中间结果能帮助多模态大语言模型(MLLM)外化空间证据和更新后的视觉状态,但其有效性取决于图像编辑器是否能准确实现所需变换。我们提出 extbf{Aphanta},一个针对MLLM→图像编辑器→MLLM流程的自动化任务发现与闭环诊断框架。Aphanta通过评估三种条件——直接推理、使用编辑器生成的中间结果、使用理想参考中间结果——来区分潜在视觉提升空间与当前编辑器的实际效用。在20个候选任务及多个编辑器-MLLM组合中,我们发现效用高度依赖任务类型:在视觉线索注入、定位和反事实状态实现方面收益明显,而对符号敏感构造或结构外推类任务则可靠性较低。在选定的正向任务子集上,集成的Qwen管道将平均任务得分从0.343提升至0.445(+10.2分;相对提升29.7%),同时完整保留被过滤和失败的任务以揭示边界。这些结果表明图像编辑应被视为专用视觉工作区而非通用推理机制,并确立Aphanta作为测量任务-表征对齐、编辑器实现能力与下游流水线效用的可复用协议。

原文摘要 · Abstract (English)

Explicit visual intermediates can help multimodal large language models (MLLMs) externalize spatial evidence and updated visual states, but their utility depends on whether an image editor can faithfully realize the required transformation. We introduce \textbf{Aphanta}, an automated task-discovery and closed-loop diagnostic framework for the MLLM -> image editor -> MLLM pipeline. Aphanta evaluates three conditions---direct reasoning, reasoning with an editor-generated intermediate, and reasoning with an idealized reference intermediate---to separate potential visual headroom from the practical utility of current editors. Across 20 candidate tasks and multiple editor--MLLM combinations, we find that utility is strongly task-conditioned. Gains concentrate in visual cue injection, grounding, and counterfactual state realization, whereas intermediates requiring symbol-sensitive construction or structural extrapolation are substantially less reliable. On the selected positive-task subset, our consolidated Qwen pipeline improves the mean task score from 0.343 to 0.445 ($+10.2$ points; $+29.7\%$ relative), while the full study also retains filtered and unsuccessful tasks to expose the boundary. These results position image editing as a specialized visual workspace rather than a universal reasoning mechanism, and establish Aphanta as a reusable protocol for measuring task--representation alignment, editor realization, and downstream pipeline utility.

多模态推理图像编辑评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。