arXiv:2511.12196cs.CVcs.HC2025-11

解决驾驶监控中视角与模态变化问题,提升跨场景识别准确率。

Cross-View Cross-Modal Unsupervised Domain Adaptation for Driver Monitoring System

  • 分两阶段学习视角不变特征并跨模态适配,无需新域标签。
  • 在Drive&Act数据集上提升50%准确率,优于现有方法5%以上。
  • 适合部署于多摄像头、多传感器车辆系统的实时监控场景。

驾驶分心是全球道路交通事故的主要原因,每年导致数千人死亡。尽管基于深度学习的驾驶行为识别方法在检测分心方面展现出潜力,但其在真实场景中的效果受限于两个关键挑战:摄像头视角差异(跨视图)和域偏移(如传感器模态或环境变化)。现有方法通常仅解决跨视图泛化或无监督域自适应之一,难以实现跨多种车型配置的鲁棒、可扩展部署。本文提出一种新颖的两阶段跨视图、跨模态无监督域自适应框架,联合处理上述问题。第一阶段在单模态内通过对比学习在多视角数据上学习视角不变且动作区分性强的特征;第二阶段利用信息瓶颈损失对新模态进行无标签域自适应。我们使用最先进的视频变换器(Video Swin、MViT)和多模态驾驶行为数据集Drive&Act进行评估,结果表明,该联合框架在RGB视频数据上的顶1准确率相比基于监督对比学习的跨视图方法提高近50%,且比仅采用无监督域自适应的方法最高提升5%,所用骨干网络相同。

原文摘要 · Abstract (English)

Driver distraction remains a leading cause of road traffic accidents, contributing to thousands of fatalities annually across the globe. While deep learning-based driver activity recognition methods have shown promise in detecting such distractions, their effectiveness in real-world deployments is hindered by two critical challenges: variations in camera viewpoints (cross-view) and domain shifts such as change in sensor modality or environment. Existing methods typically address either cross-view generalization or unsupervised domain adaptation in isolation, leaving a gap in the robust and scalable deployment of models across diverse vehicle configurations. In this work, we propose a novel two-phase cross-view, cross-modal unsupervised domain adaptation framework that addresses these challenges jointly on real-time driver monitoring data. In the first phase, we learn view-invariant and action-discriminative features within a single modality using contrastive learning on multi-view data. In the second phase, we perform domain adaptation to a new modality using information bottleneck loss without requiring any labeled data from the new domain. We evaluate our approach using state-of-the art video transformers (Video Swin, MViT) and multi modal driver activity dataset called Drive&Act, demonstrating that our joint framework improves top-1 accuracy on RGB video data by almost 50% compared to a supervised contrastive learning-based cross-view method, and outperforms unsupervised domain adaptation-only methods by up to 5%, using the same video transformer backbone.

驾驶监控域自适应多模态视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。