arXiv:2510.06243cs.CLcs.AI2025-10被引 6

用思维链结构提升多模态模型理解复杂指代的能力

CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning

  • 构建思维链式标注数据,分步解析文本指代关系
  • 在RefCOCO等数据集上提升2.5%以上准确率
  • 适合研究多模态推理与语言理解的学者参考

指代表达理解与分割是评估语言与图像融合理解能力的关键任务,常作为多模态大模型性能的基准。为此,我们提出一种新策略CoT Referring,通过结构化的思维链训练数据提升跨模态推理能力。该方法系统化地将文本结构解析为一系列指代步骤,每步识别关系并保证指代一致性,从而增强复杂查询下的准确性。我们重构了训练数据,引入新的输出格式,并为现有数据集提供新标注,构建了一个专用于复杂指代场景的评估基准。同时,将检测与分割能力整合进统一的多模态大模型框架,采用自适应加权损失进行训练。在自建基准及RefCOCO/+/g上的实验表明,该方法相比基线模型显著提升2.5%以上。

原文摘要 · Abstract (English)

Referring Expression Comprehension and Segmentation are critical tasks for assessing the integration of language understanding and image comprehension, serving as benchmarks for Multimodal Large Language Models (MLLMs) capabilities. To address these challenges, we propose a new strategy, CoT Referring, which enhances model reasoning across modalities through a structured, chain-of-thought training data structure. Our approach systematically parses textual structures to a sequential referring step, where in each step it identifies relationships and ensures consistent reference alignment, thereby improving accuracy in complex query scenarios. We restructure the training data to enforce a new output form, providing new annotations for existing datasets and compiling an evaluation benchmark from existing resources. This benchmark is designed explicitly for complex referring cases. We also integrate detection and segmentation capabilities into a unified MLLM framework, training it with a novel adaptive weighted loss to optimize performance. Experimental results on our curated benchmark and RefCOCO/+/g demonstrate the effectiveness of our approach, with a notable increase of 2.5%+ over baseline models.

多模态指代表达思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。