用视频大模型生成机器人数据摘要,提升无人值守时的监控效率
'What did the Robot do in my Absence?' Video Foundation Models to Enhance Intermittent Supervision
- 用视频大模型生成故事板、短视频和文字三种形式的摘要
- 查询驱动的摘要使信息检索准确率显著提升,40分钟数据下效果更优
- 故事板最适于物体相关查询,适合人机协同场景
本文研究视频基础模型(ViFMs)在生成机器人数据摘要中的应用,以增强对机器人团队在间歇性人类监督下的操作监控。我们提出一种新框架,能生成长时序机器人视觉数据的通用与查询驱动摘要,涵盖三类模态:故事板、短视频和文本。通过30名参与者的用户研究,评估这些摘要在40分钟无人监管期间帮助操作员准确找回观测与动作的效果。结果表明,相比通用摘要或原始数据,查询驱动摘要显著提升检索准确性,尽管任务耗时增加;故事板在物体相关查询中表现最佳。该工作据我们所知是首个零样本应用于间歇监督场景的多模态人机通信的视频大模型案例,展示了其在人机交互中的潜力与局限。
原文摘要 · Abstract (English)
This paper investigates the application of Video Foundation Models (ViFMs) for generating robot data summaries to enhance intermittent human supervision of robot teams. We propose a novel framework that produces both generic and query-driven summaries of long-duration robot vision data in three modalities: storyboards, short videos, and text. Through a user study involving 30 participants, we evaluate the efficacy of these summary methods in allowing operators to accurately retrieve the observations and actions that occurred while the robot was operating without supervision over an extended duration (40 min). Our findings reveal that query-driven summaries significantly improve retrieval accuracy compared to generic summaries or raw data, albeit with increased task duration. Storyboards are found to be the most effective presentation modality, especially for object-related queries. This work represents, to our knowledge, the first zero-shot application of ViFMs for generating multi-modal robot-to-human communication in intermittent supervision contexts, demonstrating both the promise and limitations of these models in human-robot interaction (HRI) scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。