arXiv:2601.11220cs.CL2026-01被引 1

构建多语言视觉假信息检测数据集,提升跨语言图文矛盾识别能力

MultiCaption: Detecting disinformation using multilingual visual claims

  • 构建64语言1.1万条图文对,标注是否矛盾
  • 多语言训练显著提效,无需依赖机器翻译
  • 为跨语言假信息检测提供强基准和新挑战

在线虚假信息对社会构成日益严峻的威胁,其传播速度因多媒体与多语言平台而加剧。尽管自动化事实核查技术近年取得进展,但受限于真实复杂场景下的数据稀缺。为此,我们提出MultiCaption数据集,专门用于检测视觉陈述中的矛盾。通过多种策略标注同一图像或视频的陈述对是否矛盾,最终构建包含11,088条视觉陈述、覆盖64种语言的数据集,为多模态、多语言环境下的假信息检测系统提供独特资源。我们使用基于Transformer的架构、自然语言推理模型及大语言模型进行全面实验,建立未来研究的强基线。结果表明,MultiCaption比标准自然语言推理任务更具挑战性,需任务特定微调才能获得良好表现。此外,多语言训练与测试带来的增益凸显该数据集在不依赖机器翻译的前提下构建有效多语言事实核查流程的潜力。

原文摘要 · Abstract (English)

Online disinformation poses an escalating threat to society, driven increasingly by the rapid spread of misleading content across both multimedia and multilingual platforms. While automated fact-checking methods have advanced in recent years, their effectiveness remains constrained by the scarcity of datasets that reflect these real-world complexities. To address this gap, we first present MultiCaption, a new dataset specifically designed for detecting contradictions in visual claims. Pairs of claims referring to the same image or video were labeled through multiple strategies to determine whether they contradict each other. The resulting dataset comprises 11,088 visual claims in 64 languages, offering a unique resource for building and evaluating misinformation-detection systems in truly multimodal and multilingual environments. We then provide comprehensive experiments using transformer-based architectures, natural language inference models, and large language models, establishing strong baselines for future research. The results show that MultiCaption is more challenging than standard NLI tasks, requiring task-specific finetuning for strong performance. Moreover, the gains from multilingual training and testing highlight the dataset's potential for building effective multilingual fact-checking pipelines without relying on machine translation.

假信息检测多语言视觉推理数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。