arXiv:2601.07985cs.CL2026-01

构建多语言多模态事实核查数据集,提升跨平台核查可解释性。

Multilingual, Multimodal Pipeline for Creating Authentic and Structured Fact-Checked Claim Dataset

  • 通过聚合多种来源数据并用大模型提取证据与理由,实现结构化标注。
  • 生成法语、德语双语数据集,包含文本、图像等多模态证据。
  • 适用于研究跨文化事实核查差异或开发可解释的核查模型。

在线平台虚假信息泛滥,亟需可靠、实时、可解释且多语言的事实核查资源。然而现有数据集在范围、多模态证据、结构化标注及指控、证据与结论间关联方面仍显不足。本文提出一个全面的数据收集与处理流程,通过整合ClaimReview数据源、抓取完整辟谣文章、归一化异构判断结果,并引入结构化元数据与对齐视觉内容,构建法语和德语的多模态事实核查数据集。利用先进大语言模型(LLMs)与多模态大模型,实现预定义类别下的证据提取和证据到结论的推理说明生成。通过G-Eval评估与人工评测表明,该流程支持不同机构或媒体市场间事实核查实践的细粒度比较,促进更可解释、基于证据的事实核查模型发展,并为未来多语言、多模态虚假信息验证研究奠定基础。

原文摘要 · Abstract (English)

The rapid proliferation of misinformation across online platforms underscores the urgent need for robust, up-to-date, explainable, and multilingual fact-checking resources. However, existing datasets are limited in scope, often lacking multimodal evidence, structured annotations, and detailed links between claims, evidence, and verdicts. This paper introduces a comprehensive data collection and processing pipeline that constructs multimodal fact-checking datasets in French and German languages by aggregating ClaimReview feeds, scraping full debunking articles, normalizing heterogeneous claim verdicts, and enriching them with structured metadata and aligned visual content. We used state-of-the-art large language models (LLMs) and multimodal LLMs for (i) evidence extraction under predefined evidence categories and (ii) justification generation that links evidence to verdicts. Evaluation with G-Eval and human assessment demonstrates that our pipeline enables fine-grained comparison of fact-checking practices across different organizations or media markets, facilitates the development of more interpretable and evidence-grounded fact-checking models, and lays the groundwork for future research on multilingual, multimodal misinformation verification.

事实核查多语言多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。