arXiv:2607.14683cs.AI2026-07

构建多模态车内情绪数据集,助力智能驾驶安全交互

InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring

论文配图:InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring
图 1 · 摘自论文原文
  • 融合视觉、音频与对话文本,模拟真实驾驶场景
  • 支持情绪识别、疲劳检测与分心监控三类任务
  • 适配跨语言研究,推动人车共情交互发展

理解驾驶员情绪与状态对下一代智能座舱系统至关重要。现有公开数据集多仅含视觉模态,缺乏对话信息,难以捕捉情绪背后的语言与交互线索。为此,我们提出InCarEmo,一个整合RGB与红外视频、座舱音频及对话文本的多模态数据集,基于脚本化场景模拟真实驾驶行为,覆盖多种光照条件与驾驶情境。该数据集支持三大任务:多模态情绪识别、疲劳检测与分心监控。除原始中文数据外,还构建了辅助英文基准以支持初步跨语言评估。提供统一基准与丰富基线结果,涵盖单模态与多模态方法,并分析模态缺失与噪声条件下的表现。实验表明多模态融合具显著优势,但在真实噪声与低光条件下仍存挑战。通过发布InCarEmo,旨在建立鲁棒、可解释且以人为中心的座舱情感理解基础,促进更安全、更具同理心的人车交互。

原文摘要 · Abstract (English)

Understanding driver emotion and state is critical for the next generation of intelligent in-cabin systems that ensure safety and enhance human-vehicle interaction. However, existing public datasets for in-cabin affective computing are largely limited to visual modalities and rarely include conversational information, making it difficult to capture the linguistic and interactive cues underlying driver emotion. To address these gaps, we introduce InCarEmo, a multimodal dataset for in-cabin emotion recognition and driver state monitoring. InCarEmo integrates RGB and infrared video, in-cabin audio, and dialogue text collected from scripted in-cabin scenarios designed to simulate realistic driver behaviors, covering diverse lighting conditions and driving contexts. The dataset supports three primary tasks: 1) multimodal emotion recognition, 2) fatigue detection, and 3) distraction monitoring. In addition to the original Chinese data, we construct an auxiliary English benchmark to support preliminary cross-lingual evaluation. We provide a unified benchmark with extensive baseline results across unimodal and multimodal methods, including analyses under modality-missing and noise conditions. Experimental results demonstrate the benefits of multimodal fusion and reveal remaining challenges under real-world noise and low-light conditions. By releasing InCarEmo, we aim to establish a comprehensive foundation for robust, interpretable, and human-centric in-cabin affective understanding, promoting safer and more empathetic driver-vehicle interaction.

多模态情绪识别智能座舱数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。