arXiv:2604.03156cs.CV2026-04被引 3

CAMEO通过多智能体协作实现可控图像编辑,减少结构错误。

CAMEO: A Conditional and Quality-Aware Multi-Agent Image Editing Orchestrator

论文配图:CAMEO: A Conditional and Quality-Aware Multi-Agent Image Editing Orchestrator
图 1 · 摘自论文原文
  • 分阶段多智能体架构,按需引入外部引导
  • 闭环反馈机制使编辑结果质量提升20%胜率
  • 适合需要精准控制的图像修改场景

条件图像编辑旨在根据文本提示和可选参考图修改源图像。这类编辑在需要严格结构控制的场景中至关重要(如驾驶场景中的异常插入、复杂人体姿态变换)。尽管大模型(如Seedream、Nano Banana)已有进展,但多数方法依赖单步生成,缺乏显式质量控制,易偏离原图、产生结构伪影或环境不一致修改,常需人工调优才能获得满意结果。本文提出结构化多智能体框架CAMEO,将条件编辑重构为质量感知、反馈驱动的迭代过程。CAMEO将编辑分解为规划、结构化提示、假设生成与自适应参考对齐等协同阶段,仅在任务复杂时才引入外部引导。为克服现有方法内在质量控制缺失,评估被嵌入编辑循环中,中间结果通过结构化反馈迭代优化,形成闭环修正结构与上下文不一致问题。我们在异常插入与人体姿态切换任务上评估CAMEO。在多个强基准模型及独立评估模型下,平均胜率较多个最先进模型提升20%,验证了其在可控性、鲁棒性与结构可靠性上的优势。

原文摘要 · Abstract (English)

Conditional image editing aims to modify a source image according to textual prompts and optional reference guidance. Such editing is crucial in scenarios requiring strict structural control (i.e., anomaly insertion in driving scenes and complex human pose transformation). Despite recent advances in large-scale editing models (i.e., Seedream, Nano Banana, etc), most approaches rely on single-step generation. This paradigm often lacks explicit quality control, may introduce excessive deviation from the original image, and frequently produces structural artifacts or environment-inconsistent modifications, typically requiring manual prompt tuning to achieve acceptable results. We propose \textbf{CAMEO}, a structured multi-agent framework that reformulates conditional editing as a quality-aware, feedback-driven process rather than a one-shot generation task. CAMEO decomposes editing into coordinated stages of planning, structured prompting, hypothesis generation, and adaptive reference grounding, where external guidance is invoked only when task complexity requires it. To overcome the lack of intrinsic quality control in existing methods, evaluation is embedded directly within the editing loop. Intermediate results are iteratively refined through structured feedback, forming a closed-loop process that progressively corrects structural and contextual inconsistencies. We evaluate CAMEO on anomaly insertion and human pose switching tasks. Across multiple strong editing backbones and independent evaluation models, CAMEO consistently achieves 20\% more win rate on average compared to multiple state-of-the-art models, demonstrating improved robustness, controllability, and structural reliability in conditional image editing.

图像编辑多智能体质量控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。