arXiv:2602.09084cs.CV2026-02被引 19

让图像编辑更精准:用智能体思维避免过度修改

Agent Banana: High-Fidelity Image Editing with Agentic Thinking and Tooling

  • 采用分层智能体规划执行框架,支持多轮对话式编辑
  • 在4K分辨率下实现0.871的上下文一致性与0.12的局部保真度
  • 专为专业工作流设计,适合需要高精度图像处理的场景

我们研究基于指令的图像编辑在专业工作流中的应用,发现三大挑战:(i) 编辑者常过度修改,超出用户意图;(ii) 现有模型多为单轮,多轮编辑易影响物体真实性;(iii) 评估通常在约1K分辨率进行,与实际超高清(如4K)工作流不匹配。为此提出Agent Banana,一种分层智能体规划-执行框架,支持高保真、对象感知、反思式编辑。引入两个核心机制:(1) 上下文折叠(Context Folding),将长交互历史压缩为结构化记忆,实现稳定长周期控制;(2) 图像图层分解(Image Layer Decomposition),进行局部图层级编辑,保留非目标区域,支持原生分辨率输出。为支撑严格评估,构建了HDD-Bench——一个基于对话、高分辨率的基准,包含可验证的步骤目标和原生4K图像(1180万像素),用于诊断长周期失败。在HDD-Bench上,Agent Banana在多轮一致性与背景保真度方面表现最优(如IC 0.871,SSIM-OM 0.84,LPIPS-OM 0.12),同时保持对指令遵循的竞争力,并在标准单轮编辑基准上表现优异。期望推动可靠、专业级智能体图像编辑的发展及其在真实工作流中的集成。

原文摘要 · Abstract (English)

We study instruction-based image editing under professional workflows and identify three persistent challenges: (i) editors often over-edit, modifying content beyond the user's intent; (ii) existing models are largely single-turn, while multi-turn edits can alter object faithfulness; and (iii) evaluation at around 1K resolution is misaligned with real workflows that often operate on ultra high-definition images (e.g., 4K). We propose Agent Banana, a hierarchical agentic planner-executor framework for high-fidelity, object-aware, deliberative editing. Agent Banana introduces two key mechanisms: (1) Context Folding, which compresses long interaction histories into structured memory for stable long-horizon control; and (2) Image Layer Decomposition, which performs localized layer-based edits to preserve non-target regions while enabling native-resolution outputs. To support rigorous evaluation, we build HDD-Bench, a high-definition, dialogue-based benchmark featuring verifiable stepwise targets and native 4K images (11.8M pixels) for diagnosing long-horizon failures. On HDD-Bench, Agent Banana achieves the best multi-turn consistency and background fidelity (e.g., IC 0.871, SSIM-OM 0.84, LPIPS-OM 0.12) while remaining competitive on instruction following, and also attains strong performance on standard single-turn editing benchmarks. We hope this work advances reliable, professional-grade agentic image editing and its integration into real workflows.

图像编辑智能体4K生成多轮编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。