arXiv:2608.14024cs.CV2026-08

提出跨域评估框架SSP,用相同事件比较自动驾驶视觉语言动作模型表现。

SSP: An Event-Matched Syn2Sim2Phy Cross-Domain Evaluation Framework for Autonomous Driving VLA Models

论文配图:SSP: An Event-Matched Syn2Sim2Phy Cross-Domain Evaluation Framework for Autonomous Driving VLA Models
图 1 · 摘自论文原文
  • 基于同一安全关键事件构建合成-仿真-物理三域对齐数据链
  • 物理域表现最优,但最佳模型因场景而异,得分最高0.405
  • 支持多模型输出统一评估,适合验证模型跨域鲁棒性

面向自动驾驶的视觉语言动作(VLA)模型需联合完成场景理解、语言推理与驾驶轨迹生成。现有评估常使用独立选取的合成、仿真与真实数据,导致性能差异可能由场景内容变化而非领域敏感性引起。本文提出SSP(Synthetic-Simulation-Physical)框架,通过同一安全关键交互事件实现跨域对比。从合成长尾视频出发,构建包含道路拓扑、参与方角色、相对运动、冲突演化、通行顺序、响应约束和事件阶段的事件规范。在CARLA仿真平台与封闭测试场生成平台特定实现,并在迁移审计确认关键属性保留后才进行评估。将OpenEMMA、LLaViDA和Alpamayo-R1的异构输出映射至统一语义槽与1秒轨迹窗口,评估输出有效性、语义准确性、关键交互识别、轨迹质量及风险响应能力。在切入与行人横穿两类场景中,合成、仿真、物理域的宏观平均集成VLA能力得分分别为0.259、0.291、0.325,最佳领域随场景变化;各模型得分分别为0.405、0.338、0.131。SSP提供可复现的场景迁移链与证据支撑的评估,不预设物理域绝对优越。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models for autonomous driving jointly produce scene interpretation, language-based reasoning, and driving trajectories. Existing evaluations often use independently selected synthetic, simulated, and physical data, so measured performance gaps can be confounded by changes in scenario content rather than genuine domain sensitivity. We propose SSP (Synthetic-Simulation-Physical), an event-matched Syn2Sim2Phy evaluation framework that anchors cross-domain comparison to the same safety-critical interaction. Starting from a synthetic long-tail video, SSP builds a validated event specification that preserves road topology, participant roles, relative motion, conflict evolution, passing order, response constraints, and event phases. Platform-specific realizations are then constructed in CARLA and on a closed proving ground and are evaluated only after transfer audits confirm preservation of mandatory event properties. SSP maps heterogeneous outputs from OpenEMMA, LLaViDA, and Alpamayo-R1 into common semantic slots and a 1 s trajectory window to assess output validity, semantic accuracy, critical-interaction recognition, trajectory quality, and risk response. Across Cut-in and vulnerable-road-user crossing cases, the macro-averaged Integrated VLA Capability Scores are 0.259, 0.291, and 0.325 in the Synthetic, Simulation, and Physical domains, respectively, while the best domain varies by scenario. Alpamayo-R1, OpenEMMA, and LLaViDA obtain scores of 0.405, 0.338, and 0.131. SSP provides a reproducible scene-transfer chain and an evidence-qualified evaluation of VLA behavior without assuming that the Physical domain is universally superior.

自动驾驶跨域评估VLA模型仿真测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。