arXiv:2505.17695cs.LGcs.AI2025-05被引 1

用合成数据提升复杂场景下的指代分割能力

SynRES: Towards Referring Expression Segmentation in the Wild via Synthetic Data

  • 通过密集描述生成带属性的图像-掩码-表达三元组
  • 在真实复杂场景上使模型性能提升3.8%
  • 适合研究开放域、长句指代的视觉理解任务

尽管指代表达分割(RES)基准取得进展,其评估仍局限于单一目标短查询或同一领域多目标不同查询,难以评估模型的复杂推理能力。我们提出WildRES新基准,包含长查询、多样化属性及非区分性查询,覆盖自动驾驶与机器人操作等多场景,实现对真实世界复杂推理能力的严格评估。分析发现现有模型在WildRES上性能大幅下降。为此,我们提出SynRES自动合成管道,通过三项创新生成密集配对的组合式合成数据:(1) 基于密集描述的属性丰富图像-掩码-表达三元组生成;(2) 通过图像-文本对齐分组机制修正描述与伪掩码不一致问题;(3) 采用马赛克构图与超类替换的域感知增强,强化泛化能力与属性区分性。实验表明,使用SynRES训练的模型在WildRES-ID上提升gIoU 2.0%,在WildRES-DS上提升3.8%。代码与数据集见https://github.com/UTLLab/SynRES。

原文摘要 · Abstract (English)

Despite the advances in Referring Expression Segmentation (RES) benchmarks, their evaluation protocols remain constrained, primarily focusing on either single targets with short queries (containing minimal attributes) or multiple targets from distinctly different queries on a single domain. This limitation significantly hinders the assessment of more complex reasoning capabilities in RES models. We introduce WildRES, a novel benchmark that incorporates long queries with diverse attributes and non-distinctive queries for multiple targets. This benchmark spans diverse application domains, including autonomous driving environments and robotic manipulation scenarios, thus enabling more rigorous evaluation of complex reasoning capabilities in real-world settings. Our analysis reveals that current RES models demonstrate substantial performance deterioration when evaluated on WildRES. To address this challenge, we introduce SynRES, an automated pipeline generating densely paired compositional synthetic training data through three innovations: (1) a dense caption-driven synthesis for attribute-rich image-mask-expression triplets, (2) reliable semantic alignment mechanisms rectifying caption-pseudo mask inconsistencies via Image-Text Aligned Grouping, and (3) domain-aware augmentations incorporating mosaic composition and superclass replacement to emphasize generalization ability and distinguishing attributes over object categories. Experimental results demonstrate that models trained with SynRES achieve state-of-the-art performance, improving gIoU by 2.0% on WildRES-ID and 3.8% on WildRES-DS. Code and datasets are available at https://github.com/UTLLab/SynRES.

指代分割合成数据视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。