arXiv:2509.02949cs.CLcs.CV2025-09被引 7

构建首个面向装配任务的多模态问答数据集,助力智能助手发展。

ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly

  • 采用半自动标注,结合大模型生成与人工验证,高效创建数据。
  • 包含646组多模态问答对,需理解视频与说明书的协同信息。
  • 配套81个任务图谱,适用于工业级装配助手评估与训练。

装配任务助手在日常和工业场景中潜力巨大,但相关评估资源仍不足。为推动系统发展,我们提出一个新的多模态问答数据集ProMQA-Assembly,包含646组问答对,要求理解人类活动视频与说明书的在线式协同信息。为降低成本,采用半自动化标注流程:大模型生成候选问答对,人工验证。通过引入细粒度动作标签,丰富问题类型。此外,我们构建了81个目标装配任务的任务图谱,用于基准测试及辅助人工验证。在该数据集上,我们评测了包括先进专有模型在内的多个模型,发现其包含具有挑战性的多模态问题,推理模型表现良好。我们认为该数据集将促进程序性活动助手的进一步发展。

原文摘要 · Abstract (English)

Assistants on assembly tasks show great potential to benefit humans ranging from helping with everyday tasks to interacting in industrial settings. However, evaluation resources in assembly activities are underexplored. To foster system development, we propose a new multimodal QA evaluation dataset on assembly activities. Our dataset, ProMQA-Assembly, consists of 646 QA pairs that require multimodal understanding of human activity videos and their instruction manuals in an online-style manner. For cost effectiveness in the data creation, we adopt a semi-automated QA annotation approach, where LLMs generate candidate QA pairs and humans verify them. We further improve QA generation by integrating fine-grained action labels to diversify question types. Additionally, we create 81 instruction task graphs for our target assembly tasks. These newly created task graphs are used in our benchmarking experiment, as well as in facilitating the human verification process. With our dataset, we benchmark models, including competitive proprietary multimodal models. We find that ProMQA-Assembly contains challenging multimodal questions, where reasoning models showcase promising results. We believe our new evaluation dataset contributes to the further development of procedural-activity assistants.

多模态装配任务问答数据集智能助手

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。