构建多跳多模态事实核查数据集,挑战模型综合图文证据推理能力
Piecing It All Together: Verifying Multi-Hop Multimodal Claims
- 设计多跳多模态验证任务,需跨文本、图像、表格整合证据
- 创建1.5万条复杂推理题的MMCv数据集,最新大模型仍难应对
- 适合研究多模态推理与可信AI的学者,推动可解释性验证发展
现有事实核查数据集通常不要求系统进行复杂推理或有效解读多模态证据。为解决此问题,我们提出新任务:多跳多模态事实核查。该任务要求模型在多个异构来源(包括文本、图像和表格)的证据间进行推理,判断组合证据是否支持或反驳给定主张。为此,我们构建了大型数据集MMCV,包含1.5万条多跳命题及其多模态证据,通过大语言模型生成并经人工反馈优化。实验表明,即使对最先进的多模态大模型,该任务也极具挑战性,尤其当推理步数增加时表现显著下降。此外,我们在数据集子集上建立了人类表现基准。我们希望该数据集及评估任务能推动多模态多跳事实核查的研究进展。
原文摘要 · Abstract (English)
Existing claim verification datasets often do not require systems to perform complex reasoning or effectively interpret multimodal evidence. To address this, we introduce a new task: multi-hop multimodal claim verification. This task challenges models to reason over multiple pieces of evidence from diverse sources, including text, images, and tables, and determine whether the combined multimodal evidence supports or refutes a given claim. To study this task, we construct MMCV, a large-scale dataset comprising 15k multi-hop claims paired with multimodal evidence, generated and refined using large language models, with additional input from human feedback. We show that MMCV is challenging even for the latest state-of-the-art multimodal large language models, especially as the number of reasoning hops increases. Additionally, we establish a human performance benchmark on a subset of MMCV. We hope this dataset and its evaluation task will encourage future research in multimodal multi-hop claim verification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。