arXiv:2506.13956cs.CLcs.AI2025-06被引 2

用大模型生成虚拟场景数据,提升机器人理解用户意图的能力

ASMR: Augmenting Life Scenario using Large Generative Models for Robotic Action Reflection

  • 用大语言模型模拟对话和环境,再用扩散模型生成匹配图像
  • 在真实数据集上测试,机器人动作选择准确率达最新水平
  • 适合做具身智能、人机交互的科研人员参考

设计能协助日常活动的机器人时,需结合视觉线索理解用户意图,这属于多模态分类任务。但同时包含视觉与语言的大规模数据集难以获取且耗时。为此,本文提出一种新颖的数据增强框架,聚焦于机器人辅助场景中的对话与环境图像。该方法首先利用先进的大语言模型模拟潜在对话与环境背景,再通过稳定扩散模型生成对应环境图像。生成的数据用于优化最新的多模态模型,使其在有限真实数据下更准确地判断用户交互应采取的动作。基于真实世界场景收集的数据集实验表明,该方法显著提升了机器人的动作选择能力,达到当前最优性能。

原文摘要 · Abstract (English)

When designing robots to assist in everyday human activities, it is crucial to enhance user requests with visual cues from their surroundings for improved intent understanding. This process is defined as a multimodal classification task. However, gathering a large-scale dataset encompassing both visual and linguistic elements for model training is challenging and time-consuming. To address this issue, our paper introduces a novel framework focusing on data augmentation in robotic assistance scenarios, encompassing both dialogues and related environmental imagery. This approach involves leveraging a sophisticated large language model to simulate potential conversations and environmental contexts, followed by the use of a stable diffusion model to create images depicting these environments. The additionally generated data serves to refine the latest multimodal models, enabling them to more accurately determine appropriate actions in response to user interactions with the limited target data. Our experimental results, based on a dataset collected from real-world scenarios, demonstrate that our methodology significantly enhances the robot's action selection capabilities, achieving the state-of-the-art performance.

机器人生成模型数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。