提出新基准InstructPart,让AI理解物体部件与任务指令的关联。
InstructPart: Task-Oriented Part Segmentation with Instruction Reasoning

- 构建带任务指令的物体部件分割数据集,评估模型对部件的理解能力。
- 现有视觉语言模型在部件分割任务上表现仍弱,准确率不足预期。
- 提供可直接微调的基线,性能提升一倍,适合机器人等应用研究者使用。
大型多模态基础模型在语言与视觉领域显著推进了机器人、自动驾驶、信息检索和定位等任务。然而,这些模型通常将物体视为不可分割的整体,忽视其构成部件。理解部件及其功能关联,对完成多种任务至关重要。本文提出一个全新的真实世界基准 InstructPart,包含人工标注的部件分割标签和任务导向指令,用于评估当前模型在日常场景中执行部件级任务的能力。实验表明,即使对最先进的视觉-语言模型(VLMs),任务导向的部件分割仍是挑战性问题。此外,我们提出一个简单基线,在使用本数据集微调后实现性能两倍提升。通过该数据集与基准,我们旨在推动任务导向部件分割研究,并增强VLMs在机器人、虚拟现实、信息检索等领域的适用性。
原文摘要 · Abstract (English)
Large multimodal foundation models, particularly in the domains of language and vision, have significantly advanced various tasks, including robotics, autonomous driving, information retrieval, and grounding. However, many of these models perceive objects as indivisible, overlooking the components that constitute them. Understanding these components and their associated affordances provides valuable insights into an object's functionality, which is fundamental for performing a wide range of tasks. In this work, we introduce a novel real-world benchmark, InstructPart, comprising hand-labeled part segmentation annotations and task-oriented instructions to evaluate the performance of current models in understanding and executing part-level tasks within everyday contexts. Through our experiments, we demonstrate that task-oriented part segmentation remains a challenging problem, even for state-of-the-art Vision-Language Models (VLMs). In addition to our benchmark, we introduce a simple baseline that achieves a twofold performance improvement through fine-tuning with our dataset. With our dataset and benchmark, we aim to facilitate research on task-oriented part segmentation and enhance the applicability of VLMs across various domains, including robotics, virtual reality, information retrieval, and other related fields. Project website: https://zifuwan.github.io/InstructPart/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。