HABIT用真实人类动作构建了更逼真的自动驾驶测试环境,揭示了现有系统在行人交互中的安全隐患。
HABIT: Human Action Benchmark for Interactive Traffic in CARLA
- 将真人动作数据融合进CARLA仿真,通过物理一致的动捕重定向技术生成多样行人行为
- 实测显示主流自动驾驶模型在新基准下碰撞率高达7.43次/公里,误刹率达33%
- 适合关注自动驾驶安全评估、行人交互建模的研究者和开发者使用
当前自动驾驶仿真受限于对真实人类行为的刻画不足,难以保障系统安全与可靠性。现有基准多简化行人交互,无法捕捉复杂动态意图与多样化反应。为此,我们提出HABIT(Human Action Benchmark for Interactive Traffic),一个高保真仿真基准。HABIT通过模块化、可扩展且物理一致的动作重定向流程,将来自动捕和视频的真实人类运动数据整合至CARLA(Car Learning to Act)仿真平台。从约3万条重定向动作中筛选出4,730条适配交通场景的行人动作,统一为SMPL格式以确保轨迹物理一致性。HABIT无缝集成于CARLA Leaderboard,支持自动场景生成与代理评估。安全指标如简化的伤害等级(AIS)和误刹车率(FPBR)揭示了先进自动驾驶代理此前未被发现的致命缺陷。对InterFuser、TransFuser和BEVDriver三个先进模型的评估表明:尽管在CARLA Leaderboard上碰撞率接近或等于零,但在HABIT上碰撞率最高达7.43次/公里,严重伤害风险(AIS 3+)达12.94%,且存在高达33%的非必要刹车情况。所有组件均公开发布,以支持可复现的、面向行人的智能研究。
原文摘要 · Abstract (English)
Current autonomous driving (AD) simulations are critically limited by their inadequate representation of realistic and diverse human behavior, which is essential for ensuring safety and reliability. Existing benchmarks often simplify pedestrian interactions, failing to capture complex, dynamic intentions and varied responses critical for robust system deployment. To overcome this, we introduce HABIT (Human Action Benchmark for Interactive Traffic), a high-fidelity simulation benchmark. HABIT integrates real-world human motion, sourced from mocap and videos, into CARLA (Car Learning to Act, a full autonomous driving simulator) via a modular, extensible, and physically consistent motion retargeting pipeline. From an initial pool of approximately 30,000 retargeted motions, we curate 4,730 traffic-compatible pedestrian motions, standardized in SMPL format for physically consistent trajectories. HABIT seamlessly integrates with CARLA's Leaderboard, enabling automated scenario generation and rigorous agent evaluation. Our safety metrics, including Abbreviated Injury Scale (AIS) and False Positive Braking Rate (FPBR), reveal critical failure modes in state-of-the-art AD agents missed by prior evaluations. Evaluating three state-of-the-art autonomous driving agents, InterFuser, TransFuser, and BEVDriver, demonstrates how HABIT exposes planner weaknesses that remain hidden in scripted simulations. Despite achieving close or equal to zero collisions per kilometer on the CARLA Leaderboard, the autonomous agents perform notably worse on HABIT, with up to 7.43 collisions/km and a 12.94% AIS 3+ injury risk, and they brake unnecessarily in up to 33% of cases. All components are publicly released to support reproducible, pedestrian-aware AI research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。