arXiv:2508.12842cs.CVcs.MM2025-08中稿 · ACM MM 2025 SVC Wo…被引 4

多源跨模态渐进式迁移,提升音视频骗术检测鲁棒性

Multi-source Multimodal Progressive Domain Adaption for Audio-Visual Deception Detection

  • 分阶段对齐多源音视频特征与决策,缓解域偏移问题
  • 在挑战赛第二阶段达60.43%准确率、56.99%F1分数
  • 超越第一名5.59% F1,适合多源数据骗术检测任务

本文提出针对首届视觉计算研讨会(SVC)第一届多模态骗术检测(MMDD)挑战赛的优胜方法。针对源域与目标域之间的域偏移问题,我们设计了多源多模态渐进式域自适应(MMPDA)框架,将多样源域的音视频知识迁移到目标域。通过在特征层和决策层逐步对齐源域与目标域,有效弥合了不同多模态数据集间的域差异。大量实验表明该方法有效性,最终获得第二名。在比赛第二阶段,模型取得60.43%准确率和56.99% F1得分,比第一名高出5.59% F1,比第三名高出6.75%准确率。代码已公开于 https://github.com/RH-Lin/MMPDA。

原文摘要 · Abstract (English)

This paper presents the winning approach for the 1st MultiModal Deception Detection (MMDD) Challenge at the 1st Workshop on Subtle Visual Computing (SVC). Aiming at the domain shift issue across source and target domains, we propose a Multi-source Multimodal Progressive Domain Adaptation (MMPDA) framework that transfers the audio-visual knowledge from diverse source domains to the target domain. By gradually aligning source and the target domain at both feature and decision levels, our method bridges domain shifts across diverse multimodal datasets. Extensive experiments demonstrate the effectiveness of our approach securing Top-2 place. Our approach reaches 60.43% on accuracy and 56.99\% on F1-score on competition stage 2, surpassing the 1st place team by 5.59% on F1-score and the 3rd place teams by 6.75% on accuracy. Our code is available at https://github.com/RH-Lin/MMPDA.

骗术检测域自适应多模态音视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。