arXiv:2604.16311cs.CLcs.AI2026-04

针对社交媒体图文混合信息,提出首个多模态事实核查主张提取基准与增强框架。

Multimodal Claim Extraction for Fact-Checking

  • 构建图文并茂的主张提取基准,覆盖真实场景下的多模态信息。
  • 发现现有大模型在理解修辞意图和上下文线索上表现不佳。
  • 提出感知意图的MICE框架,在关键案例中显著提升准确率。

自动事实核查依赖于主张提取作为第一步,但现有方法大多忽视了当今虚假信息的多模态特性。社交媒体帖子常结合简短非正式文本与图片(如表情包、截图、照片),这种组合带来的挑战既不同于纯文本主张提取,也不同于图像描述或视觉问答等成熟多模态任务。本文首次提出针对社交媒体的多模态主张提取基准,包含含文本与一张或多张图片的帖子,并由真实事实核查员标注出标准主张。我们采用三部分评估框架(语义对齐、忠实度、去上下文化)评估前沿多模态大模型(MLLMs),发现基线模型难以捕捉修辞意图和上下文线索。为此,我们提出MICE——一种意图感知框架,在意图敏感案例中取得显著改进。

原文摘要 · Abstract (English)

Automated Fact-Checking (AFC) relies on claim extraction as a first step, yet existing methods largely overlook the multimodal nature of today's misinformation. Social media posts often combine short, informal text with images such as memes, screenshots, and photos, creating challenges that differ from both text-only claim extraction and well-studied multimodal tasks like image captioning or visual question answering. In this work, we present the first benchmark for multimodal claim extraction from social media, consisting of posts containing text and one or more images, annotated with gold-standard claims derived from real-world fact-checkers. We evaluate state-of-the-art multimodal LLMs (MLLMs) under a three-part evaluation framework (semantic alignment, faithfulness, and decontextualization) and find that baseline MLLMs struggle to model rhetorical intent and contextual cues. To address this, we introduce MICE, an intent-aware framework which shows improvements in intent-critical cases.

多模态事实核查主张提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。