统一视觉语言模型实现精准对象分割与可控图像编辑
FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential Segmentation
- 双分支编码器同步捕捉全局语义与局部细节
- 分阶段训练使分割掩码与生成过程协同优化
- 适合需要精细图像编辑的交互式应用
近期大型视觉语言模型(LVLMs)在融合视觉理解与生成建模方面展现出潜力,支持准确的内容理解与灵活的编辑操作。然而,现有方法将‘看到什么’与‘如何编辑’分开处理:或仅进行孤立的对象分割,或仅用分割掩码作为局部编辑生成的条件提示,常依赖多个独立模型。为此,我们提出FOCUS,一个统一的LVLM框架,将感知与以对象为中心的可控生成集成于端到端结构中。FOCUS采用双分支视觉编码器,同时捕获全局语义上下文与细粒度空间细节;并利用基于MoVQGAN的视觉分词器生成离散视觉标记,提升生成质量。为实现精准可控的图像编辑,我们设计渐进式多阶段训练流程,联合优化分割掩码,并将其作为空间条件提示引导扩散解码器。该策略对齐了视觉编码、分割与生成模块,有效连接感知与细粒度合成。在三个核心任务——多模态理解、指代分割准确率、可控图像生成——上的大量实验表明,通过联合优化感知与生成能力,FOCUS实现了优异性能。
原文摘要 · Abstract (English)
Recent Large Vision Language Models (LVLMs) demonstrate promising capabilities in unifying visual understanding and generative modeling, enabling both accurate content understanding and flexible editing. However, current approaches treat "what to see" and "how to edit" separately: they either perform isolated object segmentation or utilize segmentation masks merely as conditional prompts for local edit generation tasks, often relying on multiple disjointed models. To bridge these gaps, we introduce FOCUS, a unified LVLM that integrates segmentation-aware perception and controllable object-centric generation within an end-to-end framework. FOCUS employs a dual-branch visual encoder to simultaneously capture global semantic context and fine-grained spatial details. In addition, we leverage a MoVQGAN-based visual tokenizer to produce discrete visual tokens that enhance generation quality. To enable accurate and controllable image editing, we propose a progressive multi-stage training pipeline, where segmentation masks are jointly optimized and used as spatial condition prompts to guide the diffusion decoder. This strategy aligns visual encoding, segmentation, and generation modules, effectively bridging segmentation-aware perception with fine-grained visual synthesis. Extensive experiments across three core tasks, including multimodal understanding, referring segmentation accuracy, and controllable image generation, demonstrate that FOCUS achieves strong performance by jointly optimizing visual perception and generative capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。