arXiv:2507.04061cs.CVcs.MM2025-07中稿 · ACM MM 2025被引 4

解决短视频谣言检测跨域泛化难题,提升多模态一致性与不变性。

Consistent and Invariant Generalization Learning for Short-video Misinformation Detection

  • 通过跨模态特征插值与同步蒸馏,实现视频与音频模态协同学习。
  • 引入扩散模型去噪,保留核心特征并增强跨域不变性,准确率提升12.6%。
  • 适用于需要强泛化能力的跨域谣言检测场景,尤其适合多模态数据。

短视频谣言检测在多模态领域备受关注,旨在识别带有对应音频的视频内容中的虚假信息。尽管已有显著进展,现有模型在特定领域(源域)训练后,在未见领域(目标域)常因域间差异导致性能下降。本文深入分析不同域的特性:(1) 不同域的检测主要依赖不同模态(如侧重视频或音频),因此需同时优化所有模态的表现以提升泛化能力;(2) 对于涉及跨模态联合欺诈的域,需依赖跨模态融合分析,但各模态中的域偏见(尤其是视频帧内)会在融合中累积,严重影响最终判断。为此,提出一种名为DOCTOR的新域泛化模型,包含两个关键模块:(1) 采用跨模态特征插值将多模态映射至共享空间,并通过插值蒸馏同步多模态学习;(2) 设计扩散模型添加噪声以保留多模态核心特征,利用跨模态引导去噪增强域不变特征。大量实验验证了该模型的有效性,其在多个跨域测试集上表现优异。

原文摘要 · Abstract (English)

Short-video misinformation detection has attracted wide attention in the multi-modal domain, aiming to accurately identify the misinformation in the video format accompanied by the corresponding audio. Despite significant advancements, current models in this field, trained on particular domains (source domains), often exhibit unsatisfactory performance on unseen domains (target domains) due to domain gaps. To effectively realize such domain generalization on the short-video misinformation detection task, we propose deep insights into the characteristics of different domains: (1) The detection on various domains may mainly rely on different modalities (i.e., mainly focusing on videos or audios). To enhance domain generalization, it is crucial to achieve optimal model performance on all modalities simultaneously. (2) For some domains focusing on cross-modal joint fraud, a comprehensive analysis relying on cross-modal fusion is necessary. However, domain biases located in each modality (especially in each frame of videos) will be accumulated in this fusion process, which may seriously damage the final identification of misinformation. To address these issues, we propose a new DOmain generalization model via ConsisTency and invariance learning for shORt-video misinformation detection (named DOCTOR), which contains two characteristic modules: (1) We involve the cross-modal feature interpolation to map multiple modalities into a shared space and the interpolation distillation to synchronize multi-modal learning; (2) We design the diffusion model to add noise to retain core features of multi modal and enhance domain invariant features through cross-modal guided denoising. Extensive experiments demonstrate the effectiveness of our proposed DOCTOR model. Our code is public available at https://github.com/ghh1125/DOCTOR.

谣言检测多模态域泛化扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。