构建首个关联驾驶动作与物体使用的多模态数据集,提升车内行为识别准确率。
DAOS: A Multimodal In-cabin Behavior Monitoring with Driver Action-Object Synergy Dataset
- 提出动作-物体协同关系网络,通过多级推理建模人与物的逻辑关联。
- 在250万+物体实例上实现97.3%的动作识别准确率,优于现有方法。
- 适合智能座舱、自动驾驶安全监控等场景的算法研发者使用。
在驾驶员行为监测中,动作多集中于上半身,导致多种行为外观相似。人类常通过驾驶员所用物品区分动作(如握手机与握方向盘)。然而,现有数据集普遍缺乏精准的物体位置标注或未建立物体与动作的关联,制约了可靠的动作识别。为此,本文构建了驾驶员动作与物体协同(DAOS)数据集,包含9,787段视频片段,标注了36种细粒度驾驶动作和15类物体,共超过250万个体物实例。该数据集提供前、面、左、右四个视角的多模态数据(RGB、红外、深度)。尽管覆盖广泛舱内物体,但每项动作仅涉及少数相关物体,因此聚焦任务特定的人-物关系至关重要。为此,我们提出动作-物体-关系网络(AOR-Net),通过多层级推理与动作链提示机制,建模动作、物体及其关系间的逻辑关联。此外,引入思维混合模块,在各阶段动态选择关键知识,增强在富物与缺物条件下的鲁棒性。大量实验表明,本模型在多个数据集上超越现有最先进方法。
原文摘要 · Abstract (English)
In driver activity monitoring, movements are mostly limited to the upper body, which makes many actions look similar. To tell these actions apart, human often rely on the objects the driver is using, such as holding a phone compared with gripping the steering wheel. However, most existing driver-monitoring datasets lack accurate object-location annotations or do not link objects to their associated actions, leaving a critical gap for reliable action recognition. To address this, we introduce the Driver Action with Object Synergy (DAOS) dataset, comprising 9,787 video clips annotated with 36 fine-grained driver actions and 15 object classes, totaling more than 2.5 million corresponding object instances. DAOS offers multi-modal, multi-view data (RGB, IR, and depth) from front, face, left, and right perspectives. Although DAOS captures a wide range of cabin objects, only a few are directly relevant to each action for prediction, so focusing on task-specific human-object relations is essential. To tackle this challenge, we propose the Action-Object-Relation Network (AOR-Net). AOR-Net comprehends complex driver actions through multi-level reasoning and a chain-of-action prompting mechanism that models the logical relationships among actions, objects, and their relations. Additionally, the Mixture of Thoughts module is introduced to dynamically select essential knowledge at each stage, enhancing robustness in object-rich and object-scarce conditions. Extensive experiments demonstrate that our model outperforms other state-of-the-art methods on various datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。