arXiv:2602.19359cs.ROcs.LG2026-02被引 1

用视频对比实现机器人仿真参数自动校准,还能解释原因。

Vid2Sid: Videos Can Help Close the Sim2Real Gap

  • 结合视觉大模型与优化器,通过对比仿真和真实视频诊断物理参数差异。
  • 在未见过的控制任务上表现最佳,参数误差低于13%,优于传统黑箱方法。
  • 适合需要可解释性校准的机器人研发,尤其对感知清晰场景效果突出。

校准机器人仿真中的物理参数(如摩擦、阻尼、材料刚度)通常依赖人工或黑箱优化器,难以解释误差来源。当感知仅依赖外部摄像头时,感知噪声和缺乏直接力或状态测量使问题更复杂。我们提出Vid2Sid,一种视频驱动的系统辨识流程,结合基础模型感知与视觉语言模型(VLM)闭环优化器,分析配对的仿真-真实视频,诊断具体物理不匹配,并生成带自然语言解释的参数更新建议。我们在腱驱动手指(穆贾科中的刚体动力学)和可变形连续触手(PyElastica中的软体动力学)上评估该方法。在训练中未见的测试控制任务上,Vid2Sid在所有设置中平均排名最优,性能匹配或超越黑箱优化器,且每轮迭代提供可解释推理。模拟到模拟验证表明,Vid2Sid最准确恢复真实参数(平均相对误差<13%),而黑箱方法为28%-98%。消融分析揭示三种校准范式:当感知清晰且模拟器表达能力强时,VLM引导优化表现优异;而在更复杂场景下,模型类别限制了性能。

原文摘要 · Abstract (English)

Calibrating a robot simulator's physics parameters (friction, damping, material stiffness) to match real hardware is often done by hand or with black-box optimizers that reduce error but cannot explain which physical discrepancies drive the error. When sensing is limited to external cameras, the problem is further compounded by perception noise and the absence of direct force or state measurements. We present Vid2Sid, a video-driven system identification pipeline that couples foundation-model perception with a VLM-in-the-loop optimizer that analyzes paired sim-real videos, diagnoses concrete mismatches, and proposes physics parameter updates with natural language rationales. We evaluate our approach on a tendon-actuated finger (rigid-body dynamics in MuJoCo) and a deformable continuum tentacle (soft-body dynamics in PyElastica). On sim2real holdout controls unseen during training, Vid2Sid achieves the best average rank across all settings, matching or exceeding black-box optimizers while uniquely providing interpretable reasoning at each iteration. Sim2sim validation confirms that Vid2Sid recovers ground-truth parameters most accurately (mean relative error under 13\% vs. 28--98\%), and ablation analysis reveals three calibration regimes. VLM-guided optimization excels when perception is clean and the simulator is expressive, while model-class limitations bound performance in more challenging settings.

系统辨识仿真校准视觉语言模型机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。