构建多模态流程任务问答数据集,评测模型真实场景理解能力
ProMQA: Question Answering Dataset for Multimodal Procedural Activity Understanding
- 用人类与大模型协作方式生成401组多模态流程问答对
- 当前模型在该数据集上表现远低于人类水平
- 适合评估实际应用中多模态理解能力的模型
多模态系统在指导人类完成流程性任务(如烹饪)方面具有巨大潜力。尽管应用场景多样,现有系统通常仅在动作识别或时序动作分割等传统分类任务上评估。本文提出新型评估数据集ProMQA,包含401组用户录制的流程活动(如烹饪)及其对应说明/食谱的多模态问答对。采用成本可控的人类-大模型协作标注方法:先由大模型生成问答对,再经人工验证。我们提供了基准测试结果,实验显示当前系统性能与人类存在显著差距,包括主流闭源多模态模型。希望该数据集能揭示模型多模态理解能力的新维度。
原文摘要 · Abstract (English)
Multimodal systems have great potential to assist humans in procedural activities, where people follow instructions to achieve their goals. Despite diverse application scenarios, systems are typically evaluated on traditional classification tasks, e.g., action recognition or temporal action segmentation. In this paper, we present a novel evaluation dataset, ProMQA, to measure system advancements in application-oriented scenarios. ProMQA consists of 401 multimodal procedural QA pairs on user recording of procedural activities, i.e., cooking, coupled with their corresponding instructions/recipes. For QA annotation, we take a cost-effective human-LLM collaborative approach, where the existing annotation is augmented with LLM-generated QA pairs that are later verified by humans. We then provide the benchmark results to set the baseline performance on ProMQA. Our experiment reveals a significant gap between human performance and that of current systems, including competitive proprietary multimodal models. We hope our dataset sheds light on new aspects of models' multimodal understanding capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。