arXiv:2602.01298cs.CV2026-02

让删除物体时的互动痕迹也消失,使图像更自然。

Interaction-Consistent Object Removal via MLLM-Based Reasoning

  • 用多模态大模型推理该删哪些相关元素
  • 在基准测试中比现有方法更准确
  • 适合需要精细编辑的视觉创作人群

基于图像的物体移除通常只删除指定目标,留下与之相关的交互痕迹,导致结果语义不一致。我们提出交互一致性物体移除(ICOR)问题,要求同时移除目标物体及其关联的交互元素,如依赖光照的效果、物理连接物、目标产生的元素以及情境关联物。为此,我们提出基于多模态大模型的增强型物体移除框架REORM,通过MLLM驱动分析、掩码引导移除和自校正机制实现联合删除。该框架还提供轻量本地部署版本,支持资源受限环境下的精准编辑。为评估该任务,我们构建了包含丰富交互依赖的指令驱动移除基准ICOREval。在该基准上,REORM优于当前最优图像编辑系统,证明其生成交互一致性结果的有效性。

原文摘要 · Abstract (English)

Image-based object removal often erases only the named target, leaving behind interaction evidence that renders the result semantically inconsistent. We formalize this problem as Interaction-Consistent Object Removal (ICOR), which requires removing not only the target object but also associated interaction elements, such as lighting-dependent effects, physically connected objects, targetproduced elements, and contextually linked objects. To address this task, we propose Reasoning-Enhanced Object Removal with MLLM (REORM), a reasoningenhanced object removal framework that leverages multimodal large language models to infer which elements must be jointly removed. REORM features a modular design that integrates MLLM-driven analysis, mask-guided removal, and a self-correction mechanism, along with a local-deployment variant that supports accurate editing under limited resources. To support evaluation, we introduce ICOREval, a benchmark consisting of instruction-driven removals with rich interaction dependencies. On ICOREval, REORM outperforms state-of-the-art image editing systems, demonstrating its effectiveness in producing interactionconsistent results.

图像编辑多模态一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。