构建首个含真实与合成内容的多模态可信核查数据集,助力自动识别需核查的言论。
HintsOfTruth: A Multimodal Checkworthiness Detection Dataset with Real and Synthetic Claims
- 构建27,000对图文混合的真假言论样本,涵盖真实与合成内容
- 轻量文本模型在非言论类内容识别上表现接近多模态模型
- 多模态模型更抗合成内容干扰,但计算开销大,难大规模应用
虚假信息可通过事实核查应对,但该过程成本高且耗时。识别值得核查的言论是第一步,自动化可提升效率。然而现有方法在处理多模态、跨领域及合成内容时表现不佳。本文提出 HintsOfTruth,一个公开的多模态可信核查检测数据集,包含27,000个真实世界与合成的图像/言论对。真实与合成数据混合使该数据集独特,适合评估检测方法。我们对比了微调与提示的大型语言模型(LLMs)。结果表明,配置良好的轻量级文本编码器在识别非言论类内容方面表现与多模态模型相当,但仅关注文本。多模态模型虽更准确,但计算成本高,不适用于大规模部署。面对合成数据时,多模态模型表现出更强鲁棒性。
原文摘要 · Abstract (English)
Misinformation can be countered with fact-checking, but the process is costly and slow. Identifying checkworthy claims is the first step, where automation can help scale fact-checkers' efforts. However, detection methods struggle with content that is (1) multimodal, (2) from diverse domains, and (3) synthetic. We introduce HintsOfTruth, a public dataset for multimodal checkworthiness detection with 27K real-world and synthetic image/claim pairs. The mix of real and synthetic data makes this dataset unique and ideal for benchmarking detection methods. We compare fine-tuned and prompted Large Language Models (LLMs). We find that well-configured lightweight text-based encoders perform comparably to multimodal models but the former only focus on identifying non-claim-like content. Multimodal LLMs can be more accurate but come at a significant computational cost, making them impractical for large-scale applications. When faced with synthetic data, multimodal models perform more robustly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。