用多模态机器学习检测双人互动中的说谎行为,效果优于单一模态。
Deception Detection in Dyadic Exchanges Using Multimodal Machine Learning: A Study on a Swedish Cohort
- 融合语音与面部动作、注视信息,采用晚融合策略提升识别率。
- 双人数据联合分析使准确率达71%,显著高于单人或单模态方法。
- 首次针对北欧人群研究说谎检测,适用于心理治疗等场景。
本研究探讨多模态机器学习在二人互动中检测说谎的有效性,重点关注说谎者与被欺骗者双方数据的整合。通过比较早期与晚期融合策略,利用音频和视频数据(包括面部动作单元与注视信息),在所有可能的模态与参与者组合下进行分析。数据集来自瑞典母语者在情感相关话题下的真话或谎言情境,为首次针对斯堪的纳维亚人群的研究。结果表明,结合语音与面部信息的表现优于单一模态;同时纳入双方数据显著提升检测准确率,最佳性能达71%(晚融合策略,双模态+双参与者)。该发现与心理学理论一致:初期互动中面部与发声表达受不同控制机制影响。本研究为未来在心理治疗等场景中研究双人交互奠定了基础。
原文摘要 · Abstract (English)
This study investigates the efficacy of using multimodal machine learning techniques to detect deception in dyadic interactions, focusing on the integration of data from both the deceiver and the deceived. We compare early and late fusion approaches, utilizing audio and video data - specifically, Action Units and gaze information - across all possible combinations of modalities and participants. Our dataset, newly collected from Swedish native speakers engaged in truth or lie scenarios on emotionally relevant topics, serves as the basis for our analysis. The results demonstrate that incorporating both speech and facial information yields superior performance compared to single-modality approaches. Moreover, including data from both participants significantly enhances deception detection accuracy, with the best performance (71%) achieved using a late fusion strategy applied to both modalities and participants. These findings align with psychological theories suggesting differential control of facial and vocal expressions during initial interactions. As the first study of its kind on a Scandinavian cohort, this research lays the groundwork for future investigations into dyadic interactions, particularly within psychotherapy settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。