arXiv:2510.17150cs.RO2025-10被引 2

用视觉语言模型让机器人自适应调节抓取力度,更安全地完成复杂操作。

OmniVIC: A Self-Improving Variable Impedance Controller with Vision-Language In-Context Learning for Safe Robotic Manipulation

  • 通过视觉语言模型和记忆检索,动态生成适应当前任务的阻抗参数。
  • 真实力/扭矩反馈确保接触力在安全范围内,成功率提升至61.4%。
  • 适合需要高安全性和泛化能力的复杂物理交互任务,如装配、搬运。

我们提出OmniVIC,一种由视觉语言模型(VLM)增强的通用可变阻抗控制器(VIC),旨在提升任何接触密集型机器人操作任务中的安全性和适应性。传统VIC在物理交互中表现良好,但在未见过的复杂、非结构化安全交互场景中泛化能力不足。OmniVIC通过图像与自然语言理解任务上下文,生成自适应的阻抗参数。其核心为自改进的检索增强生成(RAG)与上下文学习(ICL):RAG从结构化记忆库中检索过往相似任务经验,ICL结合这些示例与当前任务提示,调用VLM生成情境感知的阻抗参数。同时,实时力/扭矩反馈进一步调控阻抗,确保交互力在安全阈值内。实验表明,该方法在仿真与真实机器人任务中均优于基线,成功率从27%提升至61.4%,显著减少力违反。OmniVIC推动了高层语义推理与底层柔顺控制的融合,迈向更安全、更通用的操控。

原文摘要 · Abstract (English)

We present OmniVIC, a universal variable impedance controller (VIC) enhanced by a vision language model (VLM), which improves safety and adaptation in any contact-rich robotic manipulation task to enhance safe physical interaction. Traditional VIC have shown advantages when the robot physically interacts with the environment, but lack generalization in unseen, complex, and unstructured safe interactions in universal task scenarios involving contact or uncertainty. To this end, the proposed OmniVIC interprets task context derived reasoning from images and natural language and generates adaptive impedance parameters for a VIC controller. Specifically, the core of OmniVIC is a self-improving Retrieval-Augmented Generation(RAG) and in-context learning (ICL), where RAG retrieves relevant prior experiences from a structured memory bank to inform the controller about similar past tasks, and ICL leverages these retrieved examples and the prompt of current task to query the VLM for generating context-aware and adaptive impedance parameters for the current manipulation scenario. Therefore, a self-improved RAG and ICL guarantee OmniVIC works in universal task scenarios. The impedance parameter regulation is further informed by real-time force/torque feedback to ensure interaction forces remain within safe thresholds. We demonstrate that our method outperforms baselines on a suite of complex contact-rich tasks, both in simulation and on real-world robotic tasks, with improved success rates and reduced force violations. OmniVIC takes a step towards bridging high-level semantic reasoning and low-level compliant control, enabling safer and more generalizable manipulation. Overall, the average success rate increases from 27% (baseline) to 61.4% (OmniVIC).

机器人操控阻抗控制视觉语言模型安全交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。