arXiv:2412.07030cs.CLcs.AI2024-12EMNLP被引 8

用知识蒸馏生成高质量多模态多跳问答数据,提升模型性能。

FM2DS: Few-Shot Multimodal Multihop Data Synthesis with Knowledge Distillation for Question Answering

  • 五阶段流程合成多源图文问答数据,结合知识蒸馏保证质量。
  • 同等样本量下,合成数据训练模型在两个基准上平均高出1.9的准确率。
  • 首次构建长文档多模态多跳问答评测集,适合教育类复杂文本理解任务。

多模态多跳问答(MMQA)需要对来自多个来源的图像和文本进行推理。尽管视觉问答取得进展,但因缺乏高质量数据集,该多跳场景仍研究不足。现有方法多聚焦单跳、单模态或短文本,难以应对教育文档等长篇多模态内容。为此,我们提出FM2DS,首个用于生成高质量MMQA数据集的框架。其包含五阶段流程:从维基百科获取相关多模态文档,合成高层级问题与答案,并通过严格标准验证以确保质量。我们在合成数据上训练模型,并在MultimodalQA和WebQA两个基准上测试。结果表明,在相同样本量下,使用合成数据训练的模型平均比人类标注数据训练的模型在精确匹配(EM)得分上高1.9。此外,我们构建了包含1000个样本的M2QA-Bench,是首个针对长文档的MMQA基准,由FM2DS生成并经人工精修。我们认为该数据合成方法可为训练和评估MMQA模型提供坚实基础。

原文摘要 · Abstract (English)

Multimodal multihop question answering (MMQA) requires reasoning over images and text from multiple sources. Despite advances in visual question answering, this multihop setting remains underexplored due to a lack of quality datasets. Existing methods focus on single-hop, single-modality, or short texts, limiting real-world applications like interpreting educational documents with long, multimodal content. To fill this gap, we introduce FM2DS, the first framework for creating a high-quality dataset for MMQA. Our approach consists of a 5-stage pipeline that involves acquiring relevant multimodal documents from Wikipedia, synthetically generating high-level questions and answers, and validating them through rigorous criteria to ensure data quality. We evaluate our methodology by training models on our synthesized dataset and testing on two benchmarks: MultimodalQA and WebQA. Our results demonstrate that, with an equal sample size, models trained on our synthesized data outperform those trained on human-collected data by 1.9 in exact match (EM) score on average. Additionally, we introduce M2QA-Bench with 1k samples, the first benchmark for MMQA on long documents, generated using FM2DS and refined by human annotators. We believe our data synthesis method will serve as a strong foundation for training and evaluating MMQA models.

多模态问答系统数据合成知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。