arXiv:2509.04448cs.CVcs.MM2025-09EMNLP被引 15

构建可解释的多模态假信息检测模型,提升跨场景识别能力

TRUST-VL: An Explainable News Assistant for General Multimodal Misinformation Detection

  • 统一建模文本、图像与跨模态扭曲,共享推理能力
  • 在19.8万条指令数据上训练,零样本测试性能领先
  • 支持人类事实核查流程的可解释推理,适合内容审核场景

多模态假信息(包含文本、视觉及跨模态失真)因生成式AI而日益加剧社会风险。现有方法通常仅针对单一失真类型,难以泛化至未见场景。本文观察到不同失真类型共享通用推理能力,但需特定任务技能。为此提出TRUST-VL:一个统一且可解释的视觉-语言模型,用于通用多模态假信息检测。引入新型问题感知视觉增强模块,提取任务特异性视觉特征。为支持训练,构建包含19.8万样本的TRUST-Instruct指令数据集,其结构化推理链符合人类事实核查流程。在域内与零样本基准上实验表明,TRUST-VL达到当前最优性能,兼具强泛化性与可解释性。

原文摘要 · Abstract (English)

Multimodal misinformation, encompassing textual, visual, and cross-modal distortions, poses an increasing societal threat that is amplified by generative AI. Existing methods typically focus on a single type of distortion and struggle to generalize to unseen scenarios. In this work, we observe that different distortion types share common reasoning capabilities while also requiring task-specific skills. We hypothesize that joint training across distortion types facilitates knowledge sharing and enhances the model's ability to generalize. To this end, we introduce TRUST-VL, a unified and explainable vision-language model for general multimodal misinformation detection. TRUST-VL incorporates a novel Question-Aware Visual Amplifier module, designed to extract task-specific visual features. To support training, we also construct TRUST-Instruct, a large-scale instruction dataset containing 198K samples featuring structured reasoning chains aligned with human fact-checking workflows. Extensive experiments on both in-domain and zero-shot benchmarks demonstrate that TRUST-VL achieves state-of-the-art performance, while also offering strong generalization and interpretability.

假信息检测多模态可解释性视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。