arXiv:2509.24413cs.ROcs.HC2025-09被引 1

让机器人识别误导性指令,主动提醒人类避免执行错误任务

DynaMIC: Dynamic Multimodal In-Context Learning Enabled Embodied Robot Counterfactual Resistance Ability

  • 构建动态多模态上下文学习框架,识别任务中的误导指令
  • 实验验证框架能有效发现语义层面的反事实问题
  • 适合关注机器人安全与人机协作可靠性的研究者

基于自然语言的大规模预训练模型为机器人发展注入了新活力。大量研究将大模型与机器人结合,利用其强大的语义理解与生成能力,逐步实现通过自然语言指令控制机器人。然而我们发现,严格遵循人类指令(尤其是包含误导信息的指令)的机器人在执行任务时可能出错,存在安全隐患,这类似于自然语言处理中的反事实问题,在机器人研究中尚未受到足够重视。为此,本文提出了由误导性指令引发的指令反事实(DCFs),并提出DynaMIC框架,通过生成机器人任务流程来识别DCFs,并主动向人类反馈。该能力使机器人能够敏感察觉任务中的潜在反事实问题,从而提升执行过程的可靠性。我们进行了语义级实验与消融研究,验证了该框架的有效性。

原文摘要 · Abstract (English)

The emergence of large pre-trained models based on natural language has breathed new life into robotics development. Extensive research has integrated large models with robots, utilizing the powerful semantic understanding and generation capabilities of large models to facilitate robot control through natural language instructions gradually. However, we found that robots that strictly adhere to human instructions, especially those containing misleading information, may encounter errors during task execution, potentially leading to safety hazards. This resembles the concept of counterfactuals in natural language processing (NLP), which has not yet attracted much attention in robotic research. In an effort to highlight this issue for future studies, this paper introduced directive counterfactuals (DCFs) arising from misleading human directives. We present DynaMIC, a framework for generating robot task flows to identify DCFs and relay feedback to humans proactively. This capability can help robots be sensitive to potential DCFs within a task, thus enhancing the reliability of the execution process. We conducted semantic-level experiments and ablation studies, showcasing the effectiveness of this framework.

机器人人机交互反事实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。