评测大模型在多模态场景下识谎能力,发现文本模型表现优,视觉信息利用不足。
Hidden in Plain Sight: Evaluation of the Deception Detection Capabilities of LLMs in Multimodal Settings
- 对比零样本与少样本提示,用相似例选更有效
- 微调文本模型在识谎任务上达顶尖水平
- 视频动作、摘要等辅助信息提升有限,多模态融合仍弱
在日益数字化的世界中,自动识谎既关键又具挑战。本研究系统评估了大型语言模型(LLMs)和大型多模态模型(LMMs)在多个领域中的自动识谎能力。我们在三个数据集上测试:真实庭审访谈(RLTD)、人际情境下的诱导说谎(MU3D)和虚假评论(OpSpam)。对比了零样本与少样本方法,包括随机或基于相似性的上下文示例选择。结果表明,微调后的LLMs在文本识谎任务中达到当前最优性能,而LMMs未能充分利用跨模态线索。我们还分析了非语言动作、视频摘要等辅助特征的影响,以及直接输出标签与链式推理提示的有效性。研究揭示了大模型在多模态识谎中的处理机制与局限,为实际应用提供重要参考。
原文摘要 · Abstract (English)
Detecting deception in an increasingly digital world is both a critical and challenging task. In this study, we present a comprehensive evaluation of the automated deception detection capabilities of Large Language Models (LLMs) and Large Multimodal Models (LMMs) across diverse domains. We assess the performance of both open-source and commercial LLMs on three distinct datasets: real life trial interviews (RLTD), instructed deception in interpersonal scenarios (MU3D), and deceptive reviews (OpSpam). We systematically analyze the effectiveness of different experimental setups for deception detection, including zero-shot and few-shot approaches with random or similarity-based in-context example selection. Our results show that fine-tuned LLMs achieve state-of-the-art performance on textual deception detection tasks, while LMMs struggle to fully leverage cross-modal cues. Additionally, we analyze the impact of auxiliary features, such as non-verbal gestures and video summaries, and examine the effectiveness of different prompting strategies, including direct label generation and chain-of-thought reasoning. Our findings provide key insights into how LLMs process and interpret deceptive cues across modalities, highlighting their potential and limitations in real-world deception detection applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。