让视觉模型按语言指令选择并表示特定物体,实现精准可控的图像理解。
CTRL-O: Language-Controllable Object-Centric Visual Representation Learning
- 用语言描述控制每个视觉槽位的语义内容,实现对象级可调控表示
- 无需掩码监督,在复杂场景中实现物体与语言的精准绑定
- 适用于图文生成和视觉问答,支持实例级内容生成
物体中心表征学习旨在将视觉场景分解为固定大小的向量(称为“槽位”或“物体文件”),每个槽位捕捉一个独立物体。当前最先进的物体中心模型在多个领域(包括复杂真实场景)中表现出色,但在物体发现方面存在关键局限:缺乏可控性。现有模型基于预设的物体理解学习表征,无法接受用户输入来引导哪些物体被表示。将可控性引入物体中心模型可解锁一系列实用功能,如从场景中提取特定实例的表征。本文提出一种新方法,通过语言描述对槽位表征进行用户导向控制。所提出的可控物体中心表征学习方法(称为CTRL-O),可在无需掩码监督的情况下,在复杂真实场景中实现目标物体与语言的精确绑定。随后,我们将这些可控槽位表征应用于两个下游视觉-语言任务:文本到图像生成和视觉问答。该方法实现了实例级文本到图像生成,并在视觉问答任务上达到强性能。
原文摘要 · Abstract (English)
Object-centric representation learning aims to decompose visual scenes into fixed-size vectors called "slots" or "object files", where each slot captures a distinct object. Current state-of-the-art object-centric models have shown remarkable success in object discovery in diverse domains, including complex real-world scenes. However, these models suffer from a key limitation: they lack controllability. Specifically, current object-centric models learn representations based on their preconceived understanding of objects, without allowing user input to guide which objects are represented. Introducing controllability into object-centric models could unlock a range of useful capabilities, such as the ability to extract instance-specific representations from a scene. In this work, we propose a novel approach for user-directed control over slot representations by conditioning slots on language descriptions. The proposed ConTRoLlable Object-centric representation learning approach, which we term CTRL-O, achieves targeted object-language binding in complex real-world scenes without requiring mask supervision. Next, we apply these controllable slot representations on two downstream vision language tasks: text-to-image generation and visual question answering. The proposed approach enables instance-specific text-to-image generation and also achieves strong performance on visual question answering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。