arXiv:2602.01594cs.CV2026-02被引 2

统一多任务框架提升驾驶感知,缓解任务冲突导致的性能下降。

UV-M3TL: A Unified and Versatile Multimodal Multi-Task Learning Framework for Assistive Driving Perception

  • 双分支结构分离共享与特定特征,减少任务间干扰。
  • 在AIDE数据集上四项任务均达最优,跨数据集表现稳定。
  • 适合需要多模态多任务协同的智能驾驶系统开发。

高级驾驶辅助系统(ADAS)需同时理解驾驶员行为与行车环境,但联合学习异构任务易引发任务间负迁移,降低系统性能。本文提出统一且通用的多模态多任务学习框架UV-M3TL,可同步识别驾驶员行为、情绪、车辆状态及交通环境,并有效缓解任务间负迁移。框架包含两个核心组件:双分支空间通道多模态嵌入(DB-SCME)通过双分支结构显式建模任务共享与特定特征,增强跨任务知识迁移并减轻冲突;自适应特征解耦多任务损失(AFD-Loss)基于学习动态与特征解耦约束,引入自适应加权机制,提升联合优化稳定性并引导模型学习多样化表示。在AIDE数据集上的实验表明,UV-M3TL在所有四项任务上均达到当前最佳性能。为进一步验证其通用性,我们在BDD100K、CityScapes、NYUD-v2和PASCAL-Context等多个公开多任务感知基准上评估,结果表明该框架在多种任务组合中持续保持强表现,多数任务上取得最先进水平。

原文摘要 · Abstract (English)

Advanced Driver Assistance Systems (ADAS) need to understand human driver behavior while perceiving their navigation context, but jointly learning these heterogeneous tasks would cause inter-task negative transfer and impair system performance. Here, we propose a Unified and Versatile Multimodal Multi-Task Learning (UV-M3TL) framework to simultaneously recognize driver behavior, driver emotion, vehicle behavior, and traffic context, while mitigating inter-task negative transfer. Our framework incorporates two core components: dual-branch spatial channel multimodal embedding (DB-SCME) and adaptive feature-decoupled multi-task loss (AFD-Loss). DB-SCME enhances cross-task knowledge transfer while mitigating task conflicts by employing a dual-branch structure to explicitly model salient task-shared and task-specific features. AFD-Loss improves the stability of joint optimization while guiding the model to learn diverse multi-task representations by introducing an adaptive weighting mechanism based on learning dynamics and feature decoupling constraints. We evaluate our method on the AIDE dataset, and the experimental results demonstrate that UV-M3TL achieves state-of-the-art performance across all four tasks. To further prove the versatility, we evaluate UV-M3TL on additional public multi-task perception benchmarks (BDD100K, CityScapes, NYUD-v2, and PASCAL-Context), where it consistently delivers strong performance across diverse task combinations, attaining state-of-the-art results on most tasks.

多任务学习驾驶感知多模态智能驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。