arXiv:2606.15982cs.CV2026-06

研究图文编辑中模型如何发现隐含约束,揭示其失效根源。

Mind the Gap: Diagnosing Constraint Discovery Failures in Text-in-Image Editing

论文配图:Mind the Gap: Diagnosing Constraint Discovery Failures in Text-in-Image Editing
图 1 · 摘自论文原文
  • 通过文本编辑触发的约束发现任务,诊断多模态模型依赖推理能力。
  • 无引导下模型仅46%准确识别需同步修改的视觉区域,显式提示提升至94%。
  • 具体因果线索比区域名或类型标签更有效,适合关注模型可解释性的研究者。

多模态推理的关键挑战在于判断哪些视觉依赖关系在特定任务下相关,而非仅识别可见内容。本文通过图文编辑中的编辑诱导约束发现这一受控诊断场景进行研究:给定有效编辑指令与图像,模型能否识别出必须同步变化的次级区域?在461个诊断案例、四种大语言多模态模型(MLLMs)和19种约束子类型下,模型在无引导提示时仅实现46%的案例级宏观召回率,而提供明确约束时达94%,表明大量失败源于模型自身无法识别应浮现的隐含依赖。基于最优场分解的分析显示,案例特异的因果解释是效果最佳的部分引导方式(召回率0.782),优于区域名称(0.610)或类型标签(0.646),说明编辑特异性因果线索贡献了主要的性能提升。下游实验进一步表明,更高的自我发现召回率并不必然提升任务表现:未经验证的自我发现会引入误报,抵消召回增益,因而需要注重精度的约束挖掘机制。

原文摘要 · Abstract (English)

A key challenge in multimodal reasoning is determining which visual dependencies become relevant under a specific task, rather than merely recognizing visible content. We study this through edit-induced constraint discovery in text-in-image editing, a controlled diagnostic setting where a local text change can activate secondary consistency constraints: given a valid editing instruction and an image, can a model identify the secondary regions that must also change? Across 461 diagnostic cases, four MLLMs, and 19 constraint subtypes, models recover only 46% case-level macro recall under unguided prompting versus 94% when constraints are explicitly provided, suggesting that a substantial portion of the failure arises when models must decide which unstated dependencies to surface. Oracle-field decomposition shows that case-specific causal explanations are the most effective partial guidance (0.782 recall), above region names (0.610) or type labels (0.646), suggesting that edit-specific causal cues account for much of the oracle gain. A downstream experiment further shows that higher self-discovery recall does not necessarily improve task performance: unverified self-discovery introduces false positives that offset recall gains, motivating precision-aware constraint elicitation.

图文编辑多模态推理约束发现模型诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。