arXiv:2608.16556cs.AI2026-08

打通从仿真到机器人的真实执行链路,实现可追溯的物理验证与故障诊断。

DeepInsight II: One Trace from Benchmark to Robot

论文配图:DeepInsight II: One Trace from Benchmark to Robot
图 1 · 摘自论文原文
  • 统一仿真与真实机器人间的执行轨迹,保持身份一致
  • 在5个任务中验证了从仿真到物理的可复现性,修复率超80%
  • 提出可操作的故障诊断标签,适合机器人系统研发者使用

在物理AI系统中,评估成熟度与部署风险呈反比:基础模型有成熟的标准化评估体系,而实际部署依赖的具身层仍分散于各基准的仿真器、硬件和接口中。首版DeepInsight报告(v1)通过任务、资源、结果三个抽象统一了该栈的评估,但量化证据集中于基础模型层;导航与操作(系统1)及全身控制(系统0)仍为仿真案例,未涉及物理执行。DeepInsight II在此基础上固定底层架构,量化具身层表现:首先,在两个导航与四个操作基准上复现了已发布的检查点;其次,MotionBench将四个已发布全身控制器置于同一工作负载与指标契约下,实现从平行仿真到匹配机器人实验的跨域迁移,仿真与物理轨迹共享父级追踪标识,同时保留执行域特异性记录,使仿真到真实差距成为原生缩减而非工具链间调和;最后,通过系统2-1-0联合研究,将轨迹定位扩展至五个基于证据的手动切换标签,每类对应具体修复动作,并通过硬件可观测状态测试同一归因的物理验证。贡献不在于新评估架构,而在于从基准执行到匹配机器人证据的实证连续性与面向修复的诊断能力。

原文摘要 · Abstract (English)

Across a Physical AI stack, evaluation maturity is inversely aligned with deployment risk: foundation models enjoy mature, standardized harnesses, while the embodied layers on which deployment actually turns remain fragmented across benchmark-specific simulators, embodiments, and interfaces. The first DeepInsight report (v1) unified evaluation across this stack behind three abstractions---task, resource, and result---but its quantitative evidence centered on the foundation-model layer; navigation and manipulation (System 1) and whole-body control (System 0) remained simulation case studies, and physical execution was outside its empirical scope. DeepInsight II keeps that substrate fixed and quantifies the embodied half. First, it reproduces released-checkpoint references across two navigation and four manipulation benchmarks under their native protocols. Second, MotionBench places four released whole-body controllers under one workload and metric contract, then carries a qualified within-family cohort from parallel simulation to matched real-robot trials in which simulated and physical rollouts share a parent trace identity while retaining execution-domain-specific records, making the sim-to-real gap a native reduction rather than a reconciliation across toolchains. Third, a composed System 2--1--0 study extends trace localization into five evidence-grounded handoff labels, each mapped to a concrete repair action, with a measured repairability criterion and physical episodes testing the same attribution under hardware-observable state. The contribution is therefore not a new evaluation architecture, but empirical continuity from benchmark execution to matched robot evidence and repair-oriented diagnosis.

机器人仿真到现实诊断评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。