新基准评估大模型在组装家具时的细粒度时空理解能力
Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly

- 用家具组装任务设计多选题+视觉提示,测试模型对步骤顺序和部件关系的理解
- 顶尖模型在时间定位、部件匹配和空间交互上表现不佳,暴露其时序推理短板
- 适合研究视频理解、具身智能或复杂动作分析的学者使用
大型视觉语言模型(LVLMs)显著提升了视频理解能力,但现有基准主要关注粗粒度任务,如动作分类、分割、描述和检索。这些基准多依赖可口头识别的实体,难以评估真实场景下复杂的细粒度时空理解。而家具组装、烹饪等应用需要对视频进行分步精细的时空分析,当前评测体系尚未充分覆盖。为此,我们提出Flat-Pack Bench,一个聚焦家具组装任务的新基准。该基准通过多选题结合视觉提示,评估模型在装配步骤排序、状态时间定位、部件匹配与跟踪等方面的性能。实验表明,当前最优的LVLMs在细粒度时空推理方面表现显著不足,暴露出其在利用视频时序信息、追踪能力和空间交互理解上的局限。
原文摘要 · Abstract (English)
The emergence of Large Vision-Language Models (LVLMs) has significantly advanced video understanding capabilities. However, existing benchmarks focus predominantly on coarse-grained tasks such as action segmentation, classification, captioning, and retrieval. Furthermore, these benchmarks often rely on entities that can be easily identified verbally, like household objects, animals, human subjects, etc., limiting their applicability to complex, in-the-wild video scenarios. But, many applications such as furniture assembly, cooking, etc., require step-by-step fine-grained spatio-temporal understanding of the video, which is not sufficiently evaluated in current benchmarks. To address this gap, we introduce Flat-Pack Bench, a novel benchmark centered on furniture assembly tasks. Our benchmark evaluates LVLMs on nuanced tasks, including temporal ordering of assembly actions, temporal localization of assembly state, understanding part mating, and tracking, using multiple-choice questions paired with visual prompts highlighting relevant parts as references for fine-grained questions. Our experiments reveal that state-of-the-art LVLMs struggle significantly with fine-grained spatio-temporal reasoning, highlighting their limitations in effectively leveraging temporal information from videos, limited tracking ability, and understanding of spatial interactions like physical contact.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。