用多图生成机器人操作指令,自动优化评价指标提升效果
Mobile Manipulation Instruction Generation from Multiple Images with Automatic Metric Enhancement
- 结合目标物体和容器图像生成自由格式操作指令
- 新训练方法融合学习型与词频评价分数,提升指令质量
- 适合需要多视图理解的移动机器人任务研究者
针对基于目标物体图像与容器图像生成自由格式移动操作指令的问题,传统图像描述模型因架构设计偏向单图而表现不佳。本文提出一种可同时处理目标物体与容器图像的模型,生成适用于移动操作任务的自由格式指令。此外,引入一种新颖的训练方法,将基于学习和基于n-gram的自动评估指标得分作为奖励信号,使模型学会词语共现关系与合理改写。实验结果表明,该方法在标准自动评估指标上优于基线模型,包括代表性多模态大语言模型。物理实验进一步显示,使用该方法生成的指令数据增强后,可显著提升现有多模态语言理解模型在移动操作任务中的性能。
原文摘要 · Abstract (English)
We consider the problem of generating free-form mobile manipulation instructions based on a target object image and receptacle image. Conventional image captioning models are not able to generate appropriate instructions because their architectures are typically optimized for single-image. In this study, we propose a model that handles both the target object and receptacle to generate free-form instruction sentences for mobile manipulation tasks. Moreover, we introduce a novel training method that effectively incorporates the scores from both learning-based and n-gram based automatic evaluation metrics as rewards. This method enables the model to learn the co-occurrence relationships between words and appropriate paraphrases. Results demonstrate that our proposed method outperforms baseline methods including representative multimodal large language models on standard automatic evaluation metrics. Moreover, physical experiments reveal that using our method to augment data on language instructions improves the performance of an existing multimodal language understanding model for mobile manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。