让AI理解画面布局意图并精准控制生成,实现从识别到创作的统一
COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models

- 用共享专家令牌$τ_c$锚定布局意图,打通感知与生成
- 在11类布局分类上显著提升识别准确率,生成更符合提示意图
- 适合需要精确控制图像构图的研究者和设计师
构图是决定主体位置与场景组织的高层视觉意图,但现有统一多模态模型在细粒度构图识别上不可靠,且难以将此意图转化为可控生成。我们提出COMPASS,首个将构图意图控制整合于单一系统的统一多模态框架,以共享专家令牌$τ_c$作为核心意图锚点。在感知端,COMPASS以最小侵入方式将构图专长注入MoE骨干网络,并将推断出的意图蒸馏至$τ_c$;在生成端,重新使用$τ_c$作为全局条件信号,引导去噪轨迹,实现从被动分析到显式布局控制的转换。为支持大规模指令跟随式构图学习与评估,我们构建了包含11类分类体系及推理增强标注的Comp-11数据集。大量实验表明,COMPASS显著提升类别级构图理解能力,并生成更具构图一致性与提示忠实性的结果,优于强基线模型。
原文摘要 · Abstract (English)
Composition is a high-level visual intent that governs where subjects are placed and how a scene is organized, yet current unified multimodal models remain unreliable at fine-grained composition recognition and struggle to turn such intent into controllable generation. We present COMPASS, the first unified multimodal framework that grounds composition-intent control in a single system spanning both composition perception and composition-guided generation, with a shared expert token $τ_c$ as the central intent anchor. On the perception side, COMPASS injects composition expertise into an MoE backbone in a minimally invasive manner and distills the inferred intent into $τ_c$. On the generation side, COMPASS reuses $τ_c$ as a global conditioning signal that steers the denoising trajectory, effectively converting passive composition analysis into explicit layout control. To support systematic instruction-following composition learning and evaluation at scale, we construct Comp-11, a large-scale dataset with an 11-class taxonomy and reasoning-augmented annotations. Extensive experiments show that COMPASS substantially improves category-level composition understanding and delivers more composition-consistent, prompt-faithful generation than strong baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。