让AI理解自然语言指令,精准分割图像视频中的物体
Reasoning Segmentation for Images and Videos: A Survey
- 通过自然语言推理实现基于隐含语义的图像视频分割
- 系统梳理26种前沿方法与29个数据集评估体系
- 适合关注多模态理解与人机交互的研究者参考
推理分割(Reasoning Segmentation, RS)旨在根据隐含文本查询精确划分对象,其理解过程需融合推理与知识。与依赖固定语义类别或显式提示的传统分割不同,RS弥合了视觉感知与人类推理能力之间的差距,使自然语言交互更直观。本文首次全面综述图像与视频领域的推理分割研究,涵盖26项先进方法、对应评估指标,以及29个数据集与基准。我们还探讨了RS在多个领域的应用潜力,并识别现有研究空白,指明未来发展方向。
原文摘要 · Abstract (English)
Reasoning Segmentation (RS) aims to delineate objects based on implicit text queries, the interpretation of which requires reasoning and knowledge integration. Unlike the traditional formulation of segmentation problems that relies on fixed semantic categories or explicit prompting, RS bridges the gap between visual perception and human-like reasoning capabilities, facilitating more intuitive human-AI interaction through natural language. Our work presents the first comprehensive survey of RS for image and video processing, examining 26 state-of-the-art methods together with a review of the corresponding evaluation metrics, as well as 29 datasets and benchmarks. We also explore existing applications of RS across diverse domains and identify their potential extensions. Finally, we identify current research gaps and highlight promising future directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。