让大模型精准识别并操作图像中的具体物体,提升视觉理解与编辑能力。
LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation

- 基于物体中心视角,构建显式物体表示与操作机制
- 实现物体级理解、分割、编辑与生成的全流程能力
- 适合研究多模态系统精度与可控性的学者参考
大型多模态模型(LMMs)在通用视觉-语言理解任务中取得显著进展,但在需要精确物体级定位、细粒度空间推理和可控视觉操作的任务上仍受限。现有系统常难以准确定位目标实例、保持物体身份一致性,或高精度地局部化与修改特定区域。物体中心视觉提供了一种原则性框架,通过显式建模和操作视觉实体,推动多模态系统从全局场景理解迈向物体级理解、分割、编辑与生成。本文全面综述了LMMs与物体中心视觉融合的最新进展,按四大主题组织:物体中心视觉理解、物体中心指代分割、物体中心视觉编辑、物体中心视觉生成。总结了关键建模范式、学习策略与评估协议。最后讨论开放挑战与未来方向,包括鲁棒实例恒常性、细粒度空间控制、一致多步交互、统一跨任务建模及分布偏移下的可靠基准测试。期望为可扩展、精确且可信的物体中心多模态系统发展提供结构化视角。
原文摘要 · Abstract (English)
Large Multimodal Models (LMMs) have achieved remarkable progress in general-purpose vision--language understanding, yet they remain limited in tasks requiring precise object-level grounding, fine-grained spatial reasoning, and controllable visual manipulation. In particular, existing systems often struggle to identify the correct instance, preserve object identity across interactions, and localize or modify designated regions with high precision. Object-centric vision provides a principled framework for addressing these challenges by promoting explicit representations and operations over visual entities, thereby extending multimodal systems from global scene understanding to object-level understanding, segmentation, editing, and generation. This paper presents a comprehensive review of recent advances at the convergence of LMMs and object-centric vision. We organize the literature into four major themes: object-centric visual understanding, object-centric referring segmentation, object-centric visual editing, and object-centric visual generation. We further summarize the key modeling paradigms, learning strategies, and evaluation protocols that support these capabilities. Finally, we discuss open challenges and future directions, including robust instance permanence, fine-grained spatial control, consistent multi-step interaction, unified cross-task modeling, and reliable benchmarking under distribution shift. We hope this paper provides a structured perspective on the development of scalable, precise, and trustworthy object-centric multimodal systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。