arXiv:2606.00351cs.CV2026-06

无需分割掩码,精准拆解多概念图像中的目标对象

UniVerse: A Unified Modulation Framework for Segmentation-Free,Disentangled Multi-Concept Personalization

论文配图:UniVerse: A Unified Modulation Framework for Segmentation-Free,Disentangled Multi-Concept Personalization
图 1 · 摘自论文原文
  • 通过统一调制框架分解复杂场景为独立概念表征
  • 在杂乱场景中实现高精度定位与高质量视觉还原
  • 适合需要灵活个性化生成的AI艺术创作与视觉理解任务

个性化视觉理解已取得显著进展,但现有方法在输入图像包含多个物体时,难以准确定位和提取特定概念。许多先前方法依赖分割监督或表现出较差的组合泛化能力,限制了对个体概念的精确解耦与操控。本文提出UniVerse,一种用于扩散变换器的无分割、解耦式多概念个性化统一调制框架。该方法支持可组合、可分解的概念提取,可在不依赖显式分割掩码的情况下实现目标物体的细粒度定位与表征。UniVerse学习将复杂场景分解为特定概念的表示,并以统一方式组合,从而在多种视觉情境下实现鲁棒个性化。在多个基准上的大量实验表明,UniVerse在定位准确性和视觉保真度上显著优于现有最先进方法。定性与定量结果均显示,该方法能精确提取杂乱场景中的目标概念,为更灵活、可解释且个性化的视觉生成与理解开辟新路径。

原文摘要 · Abstract (English)

Personalized visual understanding has advanced significantly, yet existing approaches struggle to localize and extract specific concepts when input images contain multiple objects. Many prior methods rely heavily on segmentation-based supervision or exhibit poor compositional generalization, limiting their ability to accurately disentangle and manipulate individual concepts. In this work, we propose UniVerse, a Unified Modulation Framework for segmentation-free, disentangled multi-concept personalization in diffusion transformers. Our method allows for composable and decomposable concept extraction, enabling fine-grained localization and representation of target objects without explicit segmentation masks. UniVerse learns to decompose complex scenes into concept-specific representations and then compose them in a unified manner, enabling robust personalization across diverse visual contexts. Through extensive experiments on multiple benchmarks, we demonstrate that UniVerse significantly outperforms state-of-the-art baselines in both localization accuracy and visual fidelity. Qualitative and quantitative results show that our approach can precisely extract target concepts in cluttered scenes, paving the way for more flexible, interpretable, and personalized visual generation and understanding.

个性化生成扩散模型概念解耦无分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。