用人脸和皮肤电反应预测用户对AI的信任,提前6秒就能精准判断。
Multi-Modal Machine Learning for Early Trust Prediction in Human-AI Interaction Using Face Image and GSR Bio Signals
- 融合面部表情与皮肤电反应信号,用多模态机器学习建模信任变化。
- 早期窗口预测准确率达83%,F1-score达88%,显著优于单一模态。
- 适合医疗AI、人机交互等需实时信任感知的场景,尤其关注心理健康应用。
预测用户对AI系统的信任对安全整合基于AI的决策支持工具至关重要,尤其是在医疗领域。本研究提出一种多模态机器学习框架,结合面部视频与皮电反应(GSR)数据,在模拟注意力缺陷多动障碍(ADHD)移动健康(mHealth)情境中预测用户对AI或人类生成建议的早期信任水平。使用OpenCV提取帧,并通过预训练变换器模型进行迁移学习以获取情感特征;同时将GSR信号分解为稳态与瞬态成分以捕捉生理唤醒模式。定义了两个时间窗口:早期检测窗口(决策前6至3秒)与临近检测窗口(决策前3至0秒)。针对每个窗口分别使用图像、GSR及多模态(图像+GSR)特征进行信任预测,各模态采用机器学习算法建模,最优单模态模型通过多模态堆叠集成实现最终预测。实验表明,融合面部与生理信号显著提升预测性能:在早期窗口,准确率0.83,F1-score 0.88,ROC-AUC 0.87;在临近窗口,准确率0.75,F1-score 0.82,ROC-AUC 0.66。结果表明,生物信号可作为实时、客观的信任指标,使AI系统动态调整响应,维持合理信任水平,这对心理健康应用中避免误判诊断与治疗结果具有关键意义。
原文摘要 · Abstract (English)
Predicting human trust in AI systems is crucial for safe integration of AI-based decision support tools, especially in healthcare. This study proposes a multi-modal machine learning framework that combines image and galvanic skin response (GSR) data to predict early user trust in AI- or human-generated recommendations in a simulated ADHD mHealth context. Facial video data were processed using OpenCV for frame extraction and transferred learning with a pre-trained transformer model to derive emotional features. Concurrently, GSR signals were decomposed into tonic and phasic components to capture physiological arousal patterns. Two temporal windows were defined for trust prediction: the Early Detection Window (6 to 3 seconds before decision-making) and the Proximal Detection Window (3 to 0 seconds before decision-making). For each window, trust prediction was conducted separately using image-based, GSR-based, and multimodal (image + GSR) features. Each modality was analyzed using machine learning algorithms, and the top-performing unimodal models were integrated through a multimodal stacking ensemble for final prediction. Experimental results showed that combining facial and physiological cues significantly improved prediction performance. The multimodal stacking framework achieved an accuracy of 0.83, F1-score of 0.88, and ROC-AUC of 0.87 in the Early Detection Window, and an accuracy of 0.75, F1-score of 0.82, and ROC-AUC of 0.66 in the Proximal Detection Window. These results demonstrate the potential of bio signals as real-time, objective markers of user trust, enabling adaptive AI systems that dynamically adjust their responses to maintain calibrated trust which is a critical capability in mental health applications where mis-calibrated trust can affect diagnostic and treatment outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。