让视觉语言模型在相似物体中更准地识别,提升机器人操作的稳定性。
Confusion-Aware In-Context-Learning for Vision-Language Models in Robotic Manipulation
- 通过定位混淆源,引导模型关注易错特征。
- 在VIMA-Bench上达到85.5%成功率,泛化能力稳定。
- 适合需要高精度物体识别的机器人任务场景。
视觉语言模型(VLM)显著提升了机器人操作的泛化能力,但其在面对相似物体时仍易出现不可预测的错误。初步分析表明,这主要源于VLM固有的捷径学习问题,导致难以准确区分相似特征。为此,我们提出混淆感知上下文学习(CAICL),通过识别潜在混淆源并将其作为提示输入,引导模型聚焦于易误判特征。在VIMA-Bench上的大量实验表明,CAICL有效缓解了捷径学习问题,在不同泛化程度的任务中均表现出色,成功率达85.5%,且具备良好稳定性。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have significantly improved the generalization capabilities of robotic manipulation. However, VLM-based systems often suffer from a lack of robustness, leading to unpredictable errors, particularly in scenarios involving confusable objects. Our preliminary analysis reveals that these failures are mainly caused by shortcut learning problem inherently in VLMs, limiting their ability to accurately distinguish between confusable features. To this end, we propose Confusion-Aware In-Context Learning (CAICL), a method that enhances VLM performance in confusable scenarios for robotic manipulation. The approach begins with confusion localization and analysis, identifying potential sources of confusion. This information is then used as a prompt for the VLM to focus on features most likely to cause misidentification. Extensive experiments on the VIMA-Bench show that CAICL effectively addresses the shortcut learning issue, achieving a 85.5\% success rate and showing good stability across tasks with different degrees of generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。