arXiv:2603.20164cs.ROcs.AI2026-03中稿 · ICRA被引 1

机器人自动生成自然社交行为,靠视觉语言模型自我修正。

The Robot's Inner Critic: Self-Refinement of Social Behaviors through VLM-based Replanning

  • 用视觉语言模型当社会性‘批评家’,自动评估并优化动作
  • 在5种机器人、20个场景中评分显著优于旧方法
  • 不依赖特定接口,适配多种机器人平台

传统机器人社交行为生成受限于灵活性和自主性,依赖预设动作或人工反馈。本文提出CRISP(批判与重规划以实现交互式社交存在),一种自主框架,让机器人通过视觉语言模型(VLM)作为‘类人社会评判者’,自我批判并重规划行为。该框架整合五步:(1)分析机器人描述文件(如MJCF)提取可动关节与约束;(2)基于情境上下文生成分步行为计划;(3)参考视觉信息(关节活动范围可视化)生成底层关节控制代码;(4)利用VLM评估社会恰当性与自然度,精准定位错误步骤;(5)通过基于奖励的搜索迭代优化行为。该方法不绑定特定机器人接口,仅需结构文件即可在多种平台上生成细微差异的人类化动作。在包含五种不同机器人类型及20个场景的用户研究中,本方法在偏好度与情境适切性评分上显著优于现有方法。本研究提出一个通用框架,在最小化人工干预的同时,扩展了机器人的自主交互能力与跨平台适用性。详细结果视频与补充材料见:https://limjiyu99.github.io/inner-critic/

原文摘要 · Abstract (English)

Conventional robot social behavior generation has been limited in flexibility and autonomy, relying on predefined motions or human feedback. This study proposes CRISP (Critique-and-Replan for Interactive Social Presence), an autonomous framework where a robot critiques and replans its own actions by leveraging a Vision-Language Model (VLM) as a `human-like social critic.' CRISP integrates (1) extraction of movable joints and constraints by analyzing the robot's description file (e.g., MJCF), (2) generation of step-by-step behavior plans based on situational context, (3) generation of low-level joint control code by referencing visual information (joint range-of-motion visualizations), (4) VLM-based evaluation of social appropriateness and naturalness, including pinpointing erroneous steps, and (5) iterative refinement of behaviors through reward-based search. This approach is not tied to a specific robot API; it can generate subtly different, human-like motions on various platforms using only the robot's structure file. In a user study involving five different robot types and 20 scenarios, including mobile manipulators and humanoids, our proposed method achieved significantly higher preference and situational appropriateness ratings compared to previous methods. This research presents a general framework that minimizes human intervention while expanding the robot's autonomous interaction capabilities and cross-platform applicability. Detailed result videos and supplementary information regarding this work are available at: https://limjiyu99.github.io/inner-critic/

机器人自修正视觉语言模型社交行为

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。