用文本推理增强图像变化描述,让模型看清细微关系变化。
Leveraging Textual Compositional Reasoning for Robust Change Captioning
- 引入视觉语言模型提取场景级文本线索,补充视觉信息
- 通过三模块架构实现图文特征对齐与细粒度关系推理
- 适合需要理解复杂场景变化的应用,如医疗影像对比
变化描述旨在刻画一对图像之间的差异。然而,现有方法仅依赖视觉特征,难以捕捉细微但重要的变化,因其缺乏对对象关系和组合语义等结构化信息的显式表达能力。为此,本文提出CORTEX(COmpositional Reasoning-aware TEXt-guided)框架,通过融合互补的文本线索来增强变化理解。除像素级差异外,CORTEX利用视觉语言模型(VLMs)提供的场景级文本知识,提取揭示深层组合推理的图像-文本信号。该框架包含三个核心模块:(i) 图像级变化检测器,识别成对图像间的低层视觉差异;(ii) 推理感知文本提取(RTE)模块,利用VLM生成隐含在视觉特征中的组合推理描述;(iii) 图文双重对齐(ITDA)模块,实现视觉与文本特征的细粒度关系对齐。这使得CORTEX能够联合推理视觉与文本特征,捕捉仅靠视觉特征难以察觉的变化。
原文摘要 · Abstract (English)
Change captioning aims to describe changes between a pair of images. However, existing works rely on visual features alone, which often fail to capture subtle but meaningful changes because they lack the ability to represent explicitly structured information such as object relationships and compositional semantics. To alleviate this, we present CORTEX (COmpositional Reasoning-aware TEXt-guided), a novel framework that integrates complementary textual cues to enhance change understanding. In addition to capturing cues from pixel-level differences, CORTEX utilizes scene-level textual knowledge provided by Vision Language Models (VLMs) to extract richer image text signals that reveal underlying compositional reasoning. CORTEX consists of three key modules: (i) an Image-level Change Detector that identifies low-level visual differences between paired images, (ii) a Reasoning-aware Text Extraction (RTE) module that use VLMs to generate compositional reasoning descriptions implicit in visual features, and (iii) an Image-Text Dual Alignment (ITDA) module that aligns visual and textual features for fine-grained relational reasoning. This enables CORTEX to reason over visual and textual features and capture changes that are otherwise ambiguous in visual features alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。