构建大规模多模态人类行为数据集,助力自动驾驶安全系统研发
MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding
- 整合57000段人类动作视频与173万帧画面,覆盖多种场景来源
- 提供动作、轨迹、意图及驾驶安全关键行为的丰富标注
- 支持运动预测、生成与问答等多任务评估,适合自动驾驶研究者使用
人类是交通体系的核心组成部分,理解其行为对发展安全驾驶系统至关重要。尽管近期研究已探索人类行为的多个方面,如运动、轨迹和意图,但缺乏全面评估自动驾驶中人类行为理解能力的基准。本文提出MMHU,一个大规模多模态人类行为分析基准,包含丰富的标注信息,包括人体运动与轨迹、运动文本描述、人类意图以及与驾驶安全相关的关键行为标签。数据集涵盖57,000段人类动作片段和173万帧图像,数据来源多样,包括Waymo等主流驾驶数据集、YouTube真实场景视频及自采集数据。采用人机协同标注流程生成高质量行为描述。我们进行了详尽的数据分析,并在多个任务上进行基准测试,涵盖从运动预测到运动生成以及人类行为问答等,提供全面的评估体系。
原文摘要 · Abstract (English)
Humans are integral components of the transportation ecosystem, and understanding their behaviors is crucial to facilitating the development of safe driving systems. Although recent progress has explored various aspects of human behavior$\unicode{x2014}$such as motion, trajectories, and intention$\unicode{x2014}$a comprehensive benchmark for evaluating human behavior understanding in autonomous driving remains unavailable. In this work, we propose $\textbf{MMHU}$, a large-scale benchmark for human behavior analysis featuring rich annotations, such as human motion and trajectories, text description for human motions, human intention, and critical behavior labels relevant to driving safety. Our dataset encompasses 57k human motion clips and 1.73M frames gathered from diverse sources, including established driving datasets such as Waymo, in-the-wild videos from YouTube, and self-collected data. A human-in-the-loop annotation pipeline is developed to generate rich behavior captions. We provide a thorough dataset analysis and benchmark multiple tasks$\unicode{x2014}$ranging from motion prediction to motion generation and human behavior question answering$\unicode{x2014}$thereby offering a broad evaluation suite. Project page : https://MMHU-Benchmark.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。