arXiv:2409.14750cs.CVcs.CL2024-09EMNLP被引 25

新数据集挑战细粒度指代表达理解,检验模型精准定位与拒识能力。

FineCops-Ref: A new Dataset and Task for Fine-Grained Compositional Referring Expression Comprehension

  • 设计可调难度的细粒度指代表达数据集,覆盖多类别、属性及多跳关系
  • 引入精细编辑生成的负样本,测试模型对不可见目标的拒识能力
  • 揭示当前大模型在跨模态定位上仍有显著短板,适合研究视觉推理者

指代表达理解(REC)是一项关键的跨模态任务,客观评估语言理解、图像感知及语言到图像的对齐能力,是检验多模态大模型(MLLMs)的理想平台。为此,我们构建了一个新的REC数据集,具有两大特征:首先,数据集设计为可调控难度,要求在物体类别、属性及多跳关系上进行多层次细粒度推理;其次,通过基于现有数据的精细编辑与生成,引入负文本和负图像,以测试模型对目标对象不在图像中的情况正确拒识的能力——这一重要维度在现有数据集中常被忽略。利用该高质量数据集,我们对最先进的专用模型和多模态大模型进行了全面评估。结果表明,当前模型在实现令人满意的对齐性能方面仍存在显著差距。我们期望该数据集能激发新方法的发展,提升视觉推理能力,推动更先进的跨模态交互策略,最终释放多模态大模型的全部潜力。代码与数据集已公开于https://github.com/liujunzhuo/FineCops-Ref。

原文摘要 · Abstract (English)

Referring Expression Comprehension (REC) is a crucial cross-modal task that objectively evaluates the capabilities of language understanding, image comprehension, and language-to-image grounding. Consequently, it serves as an ideal testing ground for Multi-modal Large Language Models (MLLMs). In pursuit of this goal, we have established a new REC dataset characterized by two key features: Firstly, it is designed with controllable varying levels of difficulty, necessitating multi-level fine-grained reasoning across object categories, attributes, and multi-hop relationships. Secondly, it includes negative text and images created through fine-grained editing and generation based on existing data, thereby testing the model's ability to correctly reject scenarios where the target object is not visible in the image--an essential aspect often overlooked in existing datasets and approaches. Utilizing this high-quality dataset, we conducted comprehensive evaluations of both state-of-the-art specialist models and MLLMs. Our findings indicate that there remains a significant gap in achieving satisfactory grounding performance. We anticipate that our dataset will inspire new approaches to enhance visual reasoning and develop more advanced cross-modal interaction strategies, ultimately unlocking the full potential of MLLMs. Our code and the datasets are available at https://github.com/liujunzhuo/FineCops-Ref.

指代表达多模态视觉推理数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。