arXiv:2601.03038cs.ROcs.AI2026-01

用逻辑推理与仿真结合,验证通用机器人在复杂任务中的可靠性。

Validating Generalist Robots with Situation Calculus and STL Falsification

  • 用情境演算建模世界状态,自动生成多样任务配置
  • 在桌面操作任务中发现NVIDIA GR00T控制器的失效案例
  • 适合研发通用机器人系统的团队用于系统性测试

通用机器人正成为现实,能够理解自然语言指令并执行多样化操作。然而,其验证仍具挑战性,因每个任务对应不同的操作上下文和正确性规范,超出传统验证方法的假设范围。本文提出一种两层验证框架,结合抽象推理与具体系统伪证。在抽象层,情境演算建模世界并推导最弱先决条件,实现约束感知的组合测试,系统生成具有可控覆盖强度的多样、语义有效世界-任务配置。在具体层,这些配置被实例化用于基于仿真的伪证,结合STL监控。在桌面操作任务上的实验表明,该框架能有效发现NVIDIA GR00T控制器中的失败案例,展示了其在验证通用机器人自主性的前景。

原文摘要 · Abstract (English)

Generalist robots are becoming a reality, capable of interpreting natural language instructions and executing diverse operations. However, their validation remains challenging because each task induces its own operational context and correctness specification, exceeding the assumptions of traditional validation methods. We propose a two-layer validation framework that combines abstract reasoning with concrete system falsification. At the abstract layer, situation calculus models the world and derives weakest preconditions, enabling constraint-aware combinatorial testing to systematically generate diverse, semantically valid world-task configurations with controllable coverage strength. At the concrete layer, these configurations are instantiated for simulation-based falsification with STL monitoring. Experiments on tabletop manipulation tasks show that our framework effectively uncovers failure cases in the NVIDIA GR00T controller, demonstrating its promise for validating general-purpose robot autonomy.

机器人验证情境演算形式化方法仿真测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。