arXiv:2603.23916cs.CVcs.AI2026-03中稿 · ECCV被引 1

构建可解释的跨文化谎言检测系统,提升模型可信度与泛化能力。

DecepGPT: Schema-Driven Deception Detection with Multicultural Datasets and Robust Multimodal Learning

  • 引入结构化线索描述与推理链,生成可审计的检测报告。
  • 发布含1695样本的跨文化数据集T4-Deception,规模为现有最大。
  • 提出SICS与DMC模块,有效防止单模态过拟合,增强跨文化迁移性。

多模态谎言检测旨在通过音视频线索识别欺骗行为,服务于司法取证与安全领域。在高风险场景中,需具备可验证的证据链,将音视频线索与最终判断关联,并保证跨领域、跨文化的可靠泛化能力。然而现有基准仅提供二值标签,缺乏中间推理线索;数据集规模小、场景覆盖有限,易导致模型学习捷径。本文提出三项贡献:第一,通过扩充现有基准并添加结构化线索级描述与推理链,构建可解释的推理数据集,使模型输出可审计报告;第二,发布T4-Deception,基于全球四国统一电视节目《说真话》格式构建的跨文化数据集,共1695个样本,是目前最大的非实验室谎言检测数据集;第三,提出两个面向小样本的鲁棒学习模块:稳定个体-共性协同(SICS)通过可学习全局先验与样本自适应残差融合,结合极性感知重校准,优化多模态表示;蒸馏模态一致性(DMC)利用知识蒸馏对齐单模态预测与融合预测,防止单模态捷径学习。在三个基准及新数据集上的实验表明,该方法在域内与跨域场景下均达领先性能,且在不同文化背景间展现优异迁移能力。数据集与代码已公开。

原文摘要 · Abstract (English)

Multimodal deception detection aims to identify deceptive behavior by analyzing audiovisual cues for forensics and security. In these high-stakes settings, investigators need verifiable evidence connecting audiovisual cues to final decisions, along with reliable generalization across domains and cultural contexts. However, existing benchmarks provide only binary labels without intermediate reasoning cues. Datasets are also small with limited scenario coverage, leading to shortcut learning. We address these issues through three contributions. First, we construct reasoning datasets by augmenting existing benchmarks with structured cue-level descriptions and reasoning chains, enabling models to output auditable reports. Second, we release T4-Deception, a multicultural dataset based on the unified ``To Tell the Truth'' television format implemented across four countries. With 1695 samples, it is the largest non-laboratory deception detection dataset. Third, we propose two modules for robust learning under small-data conditions. Stabilized Individuality-Commonality Synergy (SICS) refines multimodal representations by combining learnable global priors with sample-adaptive residuals and applying polarity-aware recalibration. Distilled Modality Consistency (DMC) aligns modality-specific predictions with the fused multimodal predictions via knowledge distillation to prevent unimodal shortcut learning. Experiments on three established benchmarks and our novel dataset demonstrate that our method achieves state-of-the-art performance in both in-domain and cross-domain scenarios, while exhibiting superior transferability across diverse cultural contexts. The datasets and code are available at this link.

谎言检测多模态跨文化可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。