让机器人从视频学任务后,跨场景自动修正代码错误。
Cross-Domain Demo-to-Code via Neurosymbolic Counterfactual Reasoning
- 用符号化轨迹抽象视频动作,构建可验证的推理框架
- 在真实和模拟环境中任务成功率提升31.14%
- 适合需要跨场景迁移的机器人编程研究者
视觉语言模型(VLMs)已实现视频指导的机器人编程,使智能体能理解视频示范并生成可执行控制代码。本文将该问题建模为跨域适应问题,因示范与部署间的感知与物理差异导致程序不匹配。现有VLM缺乏对程序逻辑的理解,难以在域偏移下重构因果关系以实现任务兼容行为。为此,提出神经符号反事实推理框架NeSyCR,将视频示范抽象为捕捉底层任务流程的符号化轨迹。基于部署观测推导反事实状态,揭示跨域不相容性。通过在符号状态空间中进行可验证检查,提出程序修订以恢复与示范流程的兼容性。NeSyCR在最强基线Statler基础上实现31.14%的任务成功率提升,在仿真与真实世界操作任务中均展现出鲁棒的跨域适应能力。
原文摘要 · Abstract (English)
Recent advances in Vision-Language Models (VLMs) have enabled video-instructed robotic programming, allowing agents to interpret video demonstrations and generate executable control code. We formulate video-instructed robotic programming as a cross-domain adaptation problem, where perceptual and physical differences between demonstration and deployment induce procedural mismatches. However, current VLMs lack the procedural understanding needed to reformulate causal dependencies and achieve task-compatible behavior under such domain shifts. We introduce NeSyCR, a neurosymbolic counterfactual reasoning framework that enables verifiable adaptation of task procedures, providing a reliable synthesis of code policies. NeSyCR abstracts video demonstrations into symbolic trajectories that capture the underlying task procedure. Given deployment observations, it derives counterfactual states that reveal cross-domain incompatibilities. By exploring the symbolic state space with verifiable checks, NeSyCR proposes procedural revisions that restore compatibility with the demonstrated procedure. NeSyCR achieves a 31.14% improvement in task success over the strongest baseline Statler, showing robust cross-domain adaptation across both simulated and real-world manipulation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。