无需训练,任意大图可无缝编辑,且保持内容自然一致。
SeamEdit: A Black-Box VLM-Agnostic Pipeline for Large-Image Semantic Editing

- 将大图分块后用闭源模型做局部修复,不依赖模型内部结构
- 多阶段处理有效消除接缝痕迹与色彩错位,接缝可见度显著降低
- 适合需要高保真图像编辑的设计师或科研人员
大图像语义区域编辑需同时满足生成质量高和与周边内容自然融合。现有方法多依赖白盒模型,未充分利用闭源模型的强大生成能力;直接对拼贴图像应用闭源模型则会引发语义扭曲、画布级对齐漂移和明显接缝伪影。本文提出SeamEdit,一种免训练、模型无关的后处理流水线,将任意具备修补能力的视觉语言模型(VLM)视为黑盒代理。该流程包含五个阶段:基于叠加的分块策略、黑盒VLM修补、几何与色彩一致性校正、基于接缝风险的多候选排序,以及动态规划弯曲接缝融合。该方法有效降低接缝可见性,支持任意块区域的语义修改。
原文摘要 · Abstract (English)
Semantic region editing for large images must satisfy two requirements at the same time: high generative quality and natural integration with surrounding content. Some related methods rely on white-box models and leave the strong generation capability of closed-source models underexplored. Directly applying closed-source models to tiled editing, however, introduces several failure modes: semantic deformation, canvas-level alignment drift, and visible seam artifacts. This paper presents SeamEdit, a training-free and model-agnostic pipeline that treats any VLM with inpainting capability as a black-box oracle. SeamEdit mitigates these issues through a five-stage post-hoc pipeline: overlay-based tile decomposition, black-box VLM inpainting, geometric and color-consistency correction, seam-risk-based multi-candidate ranking, and dynamic-programming curved seam fusion. The pipeline reduces seam visibility and supports semantic modification of arbitrary tile regions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。