arXiv:2508.00265cs.CV2025-08IJCV综述被引 32

综述多模态指代分割技术,涵盖图像、视频、3D场景的精准目标定位方法。

Multimodal Referring Segmentation: A Survey

  • 统一框架归纳图像、视频、3D场景中的指代分割方法
  • 对比主流基准上的性能表现,覆盖多种视觉场景
  • 适合研究多模态理解与人机交互的学者参考

多模态指代分割旨在根据文本或音频形式的指代表达,在图像、视频和3D场景中分割出目标物体。该任务在基于用户指令进行精准物体感知的实际应用中至关重要。过去十年间,随着卷积神经网络、Transformer和大语言模型的发展,多模态感知能力显著提升,推动了该领域的广泛关注。本文全面综述多模态指代分割技术,首先介绍背景,包括问题定义和常用数据集;随后提出一个统一的元架构,并回顾图像、视频、3D场景三类主要视觉场景中的代表性方法;进一步讨论广义指代表达(GREx)方法以应对现实复杂性挑战,涵盖相关任务与实际应用;最后提供标准基准上的性能对比。相关工作持续更新于 https://github.com/henghuiding/Awesome-Multimodal-Referring-Segmentation。

原文摘要 · Abstract (English)

Multimodal referring segmentation aims to segment target objects in visual scenes, such as images, videos, and 3D scenes, based on referring expressions in text or audio format. This task plays a crucial role in practical applications requiring accurate object perception based on user instructions. Over the past decade, it has gained significant attention in the multimodal community, driven by advances in convolutional neural networks, transformers, and large language models, all of which have substantially improved multimodal perception capabilities. This paper provides a comprehensive survey of multimodal referring segmentation. We begin by introducing this field's background, including problem definitions and commonly used datasets. Next, we summarize a unified meta architecture for referring segmentation and review representative methods across three primary visual scenes, including images, videos, and 3D scenes. We further discuss Generalized Referring Expression (GREx) methods to address the challenges of real-world complexity, along with related tasks and practical applications. Extensive performance comparisons on standard benchmarks are also provided. We continually track related works at https://github.com/henghuiding/Awesome-Multimodal-Referring-Segmentation.

多模态指代分割视觉理解综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。