arXiv:2503.10225cs.CV2025-03ICCV被引 5

让模型理解被遮挡物体的完整形状并回答文本问题。

Unveiling the Invisible: Reasoning Complex Occlusions Amodally with AURA

  • 提出新任务:基于文本输入推理遮挡物体的完整形态。
  • 构建真实生活场景数据集,覆盖多样复杂遮挡情况。
  • 设计AURA模型,实现全局与空间级联合推理,支持解释性输出。

非可视分割旨在推断被遮挡物体的完整形状,即使遮挡区域外观不可见。然而,现有方法无法通过文本输入与用户交互,难以理解或推理隐含复杂的语义意图。尽管如LISA等方法将多模态大语言模型(LLMs)与分割结合用于推理,但仍仅能预测可见区域,在复杂遮挡场景中表现受限。为此,本文提出新的任务——非可视推理分割,旨在根据用户文本输入预测被遮挡物体的完整非可视形状,并提供带解释的答案。我们开发了一套可泛化的数据集生成流程,构建了一个聚焦日常场景的新数据集,涵盖多种真实世界遮挡情况。同时,提出AURA(非可视理解与推理助手)模型,采用先进全局与空间级设计,专为处理复杂遮挡而优化。大量实验证明AURA在所提数据集上的有效性。

原文摘要 · Abstract (English)

Amodal segmentation aims to infer the complete shape of occluded objects, even when the occluded region's appearance is unavailable. However, current amodal segmentation methods lack the capability to interact with users through text input and struggle to understand or reason about implicit and complex purposes. While methods like LISA integrate multi-modal large language models (LLMs) with segmentation for reasoning tasks, they are limited to predicting only visible object regions and face challenges in handling complex occlusion scenarios. To address these limitations, we propose a novel task named amodal reasoning segmentation, aiming to predict the complete amodal shape of occluded objects while providing answers with elaborations based on user text input. We develop a generalizable dataset generation pipeline and introduce a new dataset focusing on daily life scenarios, encompassing diverse real-world occlusions. Furthermore, we present AURA (Amodal Understanding and Reasoning Assistant), a novel model with advanced global and spatial-level designs specifically tailored to handle complex occlusions. Extensive experiments validate AURA's effectiveness on the proposed dataset.

图像理解视觉推理大模型分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。