用视觉语言模型扩充稀缺的人机交互数据集
ReStory: VLM-augmentation of Social Human-Robot Interaction Datasets
- 借助视觉语言模型生成可解释的互动场景故事板
- 在人工监督下实现对真实交互数据的高效扩展
- 适合人机交互研究者与交互设计者使用
人机交互(HRI)研究受限于互联网规模的数据集,因真实场景下的自然交互数据采集耗时且后勤困难。这一问题在机器人形态各异、交互方式多样时更为突出。受人机交互领域中民族志与会话分析(EMCA)研究的启发,我们提出 ReStory,一种利用视觉语言模型增强现有真实场景人机交互数据集的方法。尽管仍需人工监督,ReStory 能够生成人类可理解的互动场景故事板。我们希望该方法为 HRI 研究者和交互设计者提供一种新视角,以更有效地利用其宝贵而稀缺的数据资源。
原文摘要 · Abstract (English)
Internet-scaled datasets are a luxury for human-robot interaction (HRI) researchers, as collecting natural interaction data in the wild is time-consuming and logistically challenging. The problem is exacerbated by robots' different form factors and interaction modalities. Inspired by recent work on ethnomethodological and conversation analysis (EMCA) in the domain of HRI, we propose ReStory, a method that has the potential to augment existing in-the-wild human-robot interaction datasets leveraging Vision Language Models. While still requiring human supervision, ReStory is capable of synthesizing human-interpretable interaction scenarios in the form of storyboards. We hope our proposed approach provides HRI researchers and interaction designers with a new angle to utilizing their valuable and scarce data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。