arXiv:2409.19656cs.CL2024-09EMNLP被引 23

用合成数据训练小模型,能超越GPT-4V的谣言检测能力

Multimodal Misinformation Detection by Learning from Synthetic Data with Multimodal LLMs

论文配图:Multimodal Misinformation Detection by Learning from Synthetic Data with Multimodal LLMs
图 1 · 摘自论文原文
  • 通过两种通用数据筛选方法对齐合成与真实数据分布
  • 130亿参数小模型在真实数据集上性能提升,超过GPT-4V
  • 适合缺乏标注数据但需高精度多模态谣言检测的研究者

检测多模态虚假信息,尤其是图文组合形式,至关重要。获取大规模、高质量的真实世界事实核查数据集用于训练检测器成本高昂,因此研究者常采用AI生成的合成数据集。然而,基于合成数据训练的检测器在真实场景中的泛化能力仍不明确,主要源于数据分布差异。为此,我们提出一种从合成数据中学习的方法,通过两种模型无关的数据选择策略,实现合成数据与真实数据分布的匹配。实验表明,该方法显著提升了小型多模态大模型(130亿参数)在真实世界事实核查数据集上的表现,使其甚至超越GPT-4V。

原文摘要 · Abstract (English)

Detecting multimodal misinformation, especially in the form of image-text pairs, is crucial. Obtaining large-scale, high-quality real-world fact-checking datasets for training detectors is costly, leading researchers to use synthetic datasets generated by AI technologies. However, the generalizability of detectors trained on synthetic data to real-world scenarios remains unclear due to the distribution gap. To address this, we propose learning from synthetic data for detecting real-world multimodal misinformation through two model-agnostic data selection methods that match synthetic and real-world data distributions. Experiments show that our method enhances the performance of a small MLLM (13B) on real-world fact-checking datasets, enabling it to even surpass GPT-4V~\cite{GPT-4V}.

多模态谣言检测合成数据LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。