让照片逐步简化却不失真实感,通过智能删减与修复实现。
Progressive Photorealistic Simplification

- 用视觉语言模型识别并优先删除图像元素,逐步简化场景。
- 每一步输出都保持自然照片的逼真度与结构连贯性。
- 适合需要去杂、分层或交互编辑的真实图像处理场景。
现有图像简化技术多依赖非写实渲染(NPR),将照片转为草图、卡通或绘画风格,虽降低视觉复杂度,但牺牲了摄影真实感。本文提出一种互补方向:在保留照片真实感的前提下进行简化。我们引入渐进式语义图像简化框架,通过受控地移除并修复图像元素,逐步降低场景复杂度。每一步生成的图像仍为可信的自然照片。方法结合语义理解与生成编辑,利用视觉语言模型(VLMs)识别并优先选择需移除的元素,再通过学习到的验证器确保全程的逼真性与一致性。该过程采用迭代式‘选择-移除-验证’流程,生成高质量的简化轨迹。为进一步提升效率,我们将此过程蒸馏为端到端图像到视频生成模型,直接从单张输入图像预测连贯的简化序列。除了生成更简洁聚焦的构图外,本方法还可用于内容感知去杂、语义层分解及交互编辑。更广泛地,研究表明,在真实图像域中,通过结构化内容删除实现的简化可作为引导视觉理解的有效机制,补充传统抽象方法。
原文摘要 · Abstract (English)
Existing image simplification techniques often rely on Non-Photorealistic Rendering (NPR), transforming photographs into stylized sketches, cartoons, or paintings. While effective at reducing visual complexity, such approaches typically sacrifice photographic realism. In this work, we explore a complementary direction: simplifying images while preserving their photorealistic appearance. We introduce progressive semantic image simplification, a framework that iteratively reduces scene complexity by removing and inpainting elements in a controlled manner. At each step, the resulting image remains a plausible natural photograph. Our method combines semantic understanding with generative editing, leveraging Vision-Language Models (VLMs) to identify and prioritize elements for removal, and a learned verifier to ensure photorealism and coherence throughout the process. This is implemented via an iterative Select-Remove-Verify pipeline that produces high-quality simplification trajectories. To improve efficiency, we further distill this process into an image-to-video generation model that directly predicts coherent simplification sequences from a single input image. Beyond generating cleaner and more focused compositions, our approach enables applications such as content-aware decluttering, semantic layer decomposition, and interactive editing. More broadly, our work suggests that simplification through structured content removal can serve as a practical mechanism for guiding visual interpretation within the photorealistic domain, complementing traditional abstraction methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。