构建家具组装图文数据集,评估多模态大模型的实时辅助能力
From Instructions to Assistance: a Dataset Aligning Instruction Manuals with Assembly Videos for Evaluating Multimodal LLMs
- 构建图文对齐的家具组装数据集,支持多模态推理评估
- 模型能理解流程顺序但受限于架构与硬件,难以精准追踪步骤进展
- 适合研究多模态大模型在真实任务中的应用潜力
近期大型语言模型(LLMs)的进步推动了人工智能在复杂现实任务中的应用,促使研究从纯文本拓展至多模态场景,催生多模态大语言模型(MLMs)。随着基于LLM的助手被用于解决技术或领域特定问题,下一步是扩展其输入域,利用MLMs实现更自然的实时辅助。理想情况下,这些模型应能在用户所处环境中实时提供帮助,甚至通过虚拟现实(VR)或增强现实(AR)共享同一视角,以共同推理用户所见场景。本文旨在评估当前公开可用的MLMs在技术任务中提供此类辅助的能力。为此,我们构建了一个包含逐步标注和手册引用的家具组装数据集——手册到动作数据集(M2AD)。我们使用该数据集评估:(1)MLMs的推理能力能否减少对详细标注的依赖,从而提升标注效率并降低成本;(2)模型是否能准确追踪装配步骤进展;(3)模型是否能正确引用说明书页面。结果表明,尽管某些模型能理解流程序列,但其性能受架构和硬件限制,凸显了多图像及交错式文本-图像推理的必要性。
原文摘要 · Abstract (English)
The recent advancements introduced by Large Language Models (LLMs) have transformed how Artificial Intelligence (AI) can support complex, real world tasks, pushing research outside the text boundaries towards multi modal contexts and leading to Multimodal Large Language Models (MLMs). Given the current adoption of LLM based assistants in solving technical or domain specific problems, the natural continuation of this trend is to extend the input domains of these assistants exploiting MLMs. Ideally, these MLMs should be used as real time assistants in procedural tasks, hopefully integrating a view of the environment where the user being assisted is, or even better sharing the same point of view via Virtual Reality (VR) or Augmented Reality (AR) supports, to reason over the same scenario the user is experiencing. With this work, we aim at evaluating the quality of currently openly available MLMs to provide this kind of assistance on technical tasks. To this end, we annotated a data set of furniture assembly with step by step labels and manual references: the Manual to Action Dataset (M2AD). We used this dataset to assess (1) to which extent the reasoning abilities of MLMs can be used to reduce the need for detailed labelling, allowing for more efficient, cost effective annotation practices, (2) whether MLMs are able to track the progression of assembly steps (3) and whether MLMs can refer correctly to the instruction manual pages. Our results showed that while some models understand procedural sequences, their performance is limited by architectural and hardware constraints, highlighting the need for multi image and interleaved text image reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。