用音视频微表情提升说谎检测准确率,最高达95.4%
Enhancing Lie Detection Accuracy: A Comparative Study of Classic ML, CNN, and GCN Models using Audio-Visual Features
- 融合音频、面部微表情与手势转录,构建多模态特征
- CNN Conv1D 模型在多模态数据上达到95.4%平均准确率
- 为非侵入式测谎提供可扩展的实用框架,适合安全与司法场景
测谎仪误判常导致冤案、虚假信息及偏见,对法律与政治体系造成严重影响。近年来,分析面部微表情成为测谎新方法,但现有模型准确率与泛化能力仍不足。本研究旨在改善此问题。通过结合听觉输入、视觉面部微表情及人工转录的手势标注,构建独特的多模态变压器架构,推动更可靠的非侵入式测谎模型发展。使用Vision Transformer提取视觉特征,OpenSmile提取音频特征,并与参与者微表情和手势的转录文本拼接。在这些处理后的多模态特征上训练多种分类模型,用于区分谎言与真相。其中,CNN Conv1D 多模态模型实现95.4%的平均准确率。但要实现更高品质数据集与更强泛化能力,仍需进一步研究。
原文摘要 · Abstract (English)
Inaccuracies in polygraph tests often lead to wrongful convictions, false information, and bias, all of which have significant consequences for both legal and political systems. Recently, analyzing facial micro-expressions has emerged as a method for detecting deception; however, current models have not reached high accuracy and generalizability. The purpose of this study is to aid in remedying these problems. The unique multimodal transformer architecture used in this study improves upon previous approaches by using auditory inputs, visual facial micro-expressions, and manually transcribed gesture annotations, moving closer to a reliable non-invasive lie detection model. Visual and auditory features were extracted using the Vision Transformer and OpenSmile models respectively, which were then concatenated with the transcriptions of participants micro-expressions and gestures. Various models were trained for the classification of lies and truths using these processed and concatenated features. The CNN Conv1D multimodal model achieved an average accuracy of 95.4%. However, further research is still required to create higher-quality datasets and even more generalized models for more diverse applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。