用数字孪生表征分离视觉感知与文本推理,提升图像指代分割性能。
Decoupling the Image Perception and Multimodal Reasoning for Reasoning Segmentation with Digital Twin Representations
- 将图像转为数字孪生表示,保留物体空间关系
- 使用大语言模型在孪生表示上进行显式推理,准确率领先
- 适合需要复杂多模态推理的视觉任务研究者
推理分割(RS)是要求根据隐含文本查询分割目标对象的多模态视觉任务,需兼具精确视觉感知与跨模态推理能力。现有方法依赖微调视觉语言模型同时完成感知与推理,但图像分词破坏了物体间的连续空间关系。本文提出DTwinSeger,利用数字孪生(DT)表示作为中间层,解耦感知与推理过程。创新性地将RS重构为两阶段:第一阶段将图像转换为保持空间结构与语义特性的结构化DT表示;第二阶段由大语言模型(LLM)在此表示上执行显式推理以定位目标对象。我们还提出了针对DT表示的监督微调方法及配套数据集Seg-DT,增强LLM对DT表示的推理能力。实验表明,该方法在两个图像推理分割基准和三个图像指代分割基准上均达到当前最优性能,证明了DT表示作为视觉与文本间有效桥梁的作用,仅用一个LLM即可完成复杂多模态推理任务。
原文摘要 · Abstract (English)
Reasoning Segmentation (RS) is a multimodal vision-text task that requires segmenting objects based on implicit text queries, demanding both precise visual perception and vision-text reasoning capabilities. Current RS approaches rely on fine-tuning vision-language models (VLMs) for both perception and reasoning, but their tokenization of images fundamentally disrupts continuous spatial relationships between objects. We introduce DTwinSeger, a novel RS approach that leverages Digital Twin (DT) representation as an intermediate layer to decouple perception from reasoning. Innovatively, DTwinSeger reformulates RS as a two-stage process, where the first transforms the image into a structured DT representation that preserves spatial relationships and semantic properties and then employs a Large Language Model (LLM) to perform explicit reasoning over this representation to identify target objects. We propose a supervised fine-tuning method specifically for LLM with DT representation, together with a corresponding fine-tuning dataset Seg-DT, to enhance the LLM's reasoning capabilities with DT representations. Experiments show that our method can achieve state-of-the-art performance on two image RS benchmarks and three image referring segmentation benchmarks. It yields that DT representation functions as an effective bridge between vision and text, enabling complex multimodal reasoning tasks to be accomplished solely with an LLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。