arXiv:2604.05947cs.CV2026-04

让不同模态自动协作,精准识别司机动作

Mixture-of-Modality-Experts with Holistic Token Learning for Fine-Grained Multimodal Visual Analytics in Driver Action Recognition

  • 按模态分组专家,动态选择可靠信息
  • 通过全局与时空令牌提升细节理解能力
  • 适合需要精细多模态分析的自动驾驶场景

当异构模态提供互补但依赖输入的证据时,鲁棒的多模态视觉分析仍具挑战。现有方法多依赖固定融合模块或预设跨模态交互,难以适应模态可靠性变化并捕捉细粒度动作线索。为此,我们提出混合模态专家(MoME)框架与整体令牌学习(HTL)策略。MoME实现模态专用专家间的自适应协作,而HTL通过类别令牌与时空令牌提升专家内精炼与专家间知识迁移。该方法构建以知识为中心的多模态学习框架,增强专家专属性,降低融合模糊性。我们在司机动作识别任务上验证该框架,公开基准测试结果表明,所提MoME框架与HTL策略联合优于代表性单模态及多模态基线。额外消融、验证与可视化结果进一步证实,HTL策略提升了细微多模态理解能力,并增强了可解释性。

原文摘要 · Abstract (English)

Robust multimodal visual analytics remains challenging when heterogeneous modalities provide complementary but input-dependent evidence for decision-making.Existing multimodal learning methods mainly rely on fixed fusion modules or predefined cross-modal interactions, which are often insufficient to adapt to changing modality reliability and to capture fine-grained action cues. To address this issue, we propose a Mixture-of-Modality-Experts (MoME) framework with a Holistic Token Learning (HTL) strategy. MoME enables adaptive collaboration among modality-specific experts, while HTL improves both intra-expert refinement and inter-expert knowledge transfer through class tokens and spatio-temporal tokens. In this way, our method forms a knowledge-centric multimodal learning framework that improves expert specialization while reducing ambiguity in multimodal fusion.We validate the proposed framework on driver action recognition as a representative multimodal understanding taskThe experimental results on the public benchmark show that the proposed MoME framework and the HTL strategy jointly outperform representative single-modal and multimodal baselines. Additional ablation, validation, and visualization results further verify that the proposed HTL strategy improves subtle multimodal understanding and offers better interpretability.

多模态学习司机行为识别专家模型视觉分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。