arXiv:2606.18988cs.AI2026-06

用可解释的推理框架检测多模态欺骗行为,提升准确率与透明度。

ThinkDeception: A Progressive Reinforcement Learning Framework for Interpretable Multimodal Deception Detection

论文配图:ThinkDeception: A Progressive Reinforcement Learning Framework for Interpretable Multimodal Deception Detection
图 1 · 摘自论文原文
  • 引入多模态大模型,将欺骗检测转为逐步推理过程。
  • 在主流数据集上达到新SOTA,准确率显著优于现有方法。
  • 设计渐进式训练策略,增强模型对跨模态不一致的捕捉能力。

多模态欺骗检测对识别欺诈意图至关重要,但现有方法多依赖黑箱端到端范式,缺乏可解释性,难以提供透明的推理路径,也难以显式捕捉欺骗行为中隐含的细微跨模态不一致性。为此,我们提出ThinkDeception,一个新颖且可解释的多模态欺骗检测框架。作为开创性工作,首次将多模态大语言模型(MLLMs)引入该领域,将欺骗检测从传统二分类任务转变为明确的认知推理过程。基于首个精心标注的多模态思维链(CoT)数据集,我们构建了基础模型ThinkDeception Base,实证验证了模态不一致性在识破欺骗中的关键作用。在此基础上,核心创新在于提出视觉-音频一致性组相对策略优化(VAC-GRPO),并采用渐进式训练策略。不同于标准GRPO,我们将训练数据分为四个递进难度层级,引导模型经历心理上合理由易到难的认知过渡。通过创新性地结合动态课程调度器、多维度过程感知奖励机制与反思学习范式,显著提升了模型的整体推理质量。在主流基准上的大量实验表明,ThinkDeception实现了新的最优性能,大幅超越现有方法在检测准确率与推理质量上的表现。最终,本工作成功推动欺骗检测向可解释的多模态认知推理方向发展。

原文摘要 · Abstract (English)

Multimodal deception detection is critical for identifying fraudulent intentions, yet existing approaches predominantly rely on end to end black--box paradigms. These methods suffer from a severe lack of interpretability failing to provide transparent reasoning trajectories and struggling to explicitly capture the subtle, cross modal inconsistencies inherent in deceptive behaviors. To transcend these limitations, we propose ThinkDeception, a novel and interpretable multimodal deception detection framework. As a pioneering effort, it introduces Multimodal Large Language Models (MLLMs) into this domain, transforming deception detection from a traditional binary classification task into an explicit cognitive reasoning process. Facilitated by the first meticulously annotated step--by--step multimodal Chain of Thought (CoT) dataset, we develop a foundational model, ThinkDeception Base, empirically validating the critical role of modal inconsistency in decoding deception. Building upon this foundation, our core innovation lies in proposing Visual-Audio Consistency Group Relative Policy Optimization(VAC--GRPO) equipped with a progressive training strategy. Distinct from standard GRPO, we stratify the training data into four progressive difficulty tiers, guiding the model through a psychologically grounded easy--to--hard cognitive transition. By innovatively coupling this dynamic curriculum scheduler with a multi dimensional, process aware reward mechanism and a reflective learning paradigm, we significantly elevate the model's overall reasoning quality. Extensive experiments on mainstream benchmarks demonstrate that ThinkDeception establishes a new SOTA, significantly outperforming existing methods in both detection accuracy and rationale quality. Ultimately, this work successfully drives the field of deception detection toward interpretable, multimodal cognitive reasoning.

欺骗检测可解释性多模态强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。