构建日语人机对话数据集,提升机器人对物体属性与动作的感知能力
J-ORA: A Framework and Multimodal Dataset for Japanese Object Identification, Reference, Action Prediction in Robot Perception

- 基于日语对话场景,标注物体类别、颜色、形状等多维属性
- 引入属性信息后,多模态模型性能显著提升,但开源与闭源模型仍有差距
- 揭示不同模型在理解物体功能和上下文关系上的差异,适合机器人交互研究者
我们提出 J-ORA,一个新型多模态数据集,填补了日本语境下人机对话中机器人感知的空白。该数据集通过详尽的物体属性标注(如类别、颜色、形状、大小、材质、空间关系),支持物体识别、指代消解和下一步动作预测三项关键任务。对自研及开源视觉语言模型(VLMs)的广泛评估表明,引入详细物体属性可显著提升多模态感知性能。尽管如此,仍存在闭源与开源 VLM 间的性能差距。对物体功能性的分析显示,不同 VLM 在理解物体用途和上下文关系方面表现各异。这些发现凸显了丰富、上下文敏感的属性标注在动态环境中推进机器人感知的重要性。
原文摘要 · Abstract (English)
We introduce J-ORA, a novel multimodal dataset that bridges the gap in robot perception by providing detailed object attribute annotations within Japanese human-robot dialogue scenarios. J-ORA is designed to support three critical perception tasks, object identification, reference resolution, and next-action prediction, by leveraging a comprehensive template of attributes (e.g., category, color, shape, size, material, and spatial relations). Extensive evaluations with both proprietary and open-source Vision Language Models (VLMs) reveal that incorporating detailed object attributes substantially improves multimodal perception performance compared to without object attributes. Despite the improvement, we find that there still exists a gap between proprietary and open-source VLMs. In addition, our analysis of object affordances demonstrates varying abilities in understanding object functionality and contextual relationships across different VLMs. These findings underscore the importance of rich, context-sensitive attribute annotations in advancing robot perception in dynamic environments. See project page at https://jatuhurrra.github.io/J-ORA/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。