PRISM数据集提升AI在零售场景下的空间与动作理解能力
PRISM: A Multi-View Multi-Capability Retail Video Dataset for Embodied Vision-Language Models
- 构建三维知识体系,覆盖空间、物理与行动三类认知
- 270万样本,1180万帧视频,细粒度标注支持多视角学习
- 微调后动作理解准确率提升36.4%,显著改善真实场景表现
当前先进物理AI模型在通用视觉理解上表现优异,但在结构化现实部署环境中仍面临感知挑战。我们提出PRISM,一个面向真实零售场景的多视角、多能力视频监督微调(SFT)语料库,包含270万样本。该数据集基于全新三维知识本体,涵盖空间、时间与物理、具身动作三类知识,覆盖20多个能力探针,评估维度包括具身推理(ER)、常识(CS)、空间感知(SP)和直觉物理(IP)。PRISM是首个在同一真实应用场景中完整实现三类知识标注的数据集。数据采集自五个超市的内视、外视及360°视角,包含开放式、链式思考与多选题标注形式。视频以4帧/秒速率,总计约1180万帧与7.3亿个标记,属于目前最大的领域特定视频SFT语料库之一。在该数据集上微调,使所有20+探针的错误率相比预训练基线降低66.6%,尤其在具身动作理解方面,准确率提升达36.4%。结果表明,基于本体结构化的领域专用微调能有效增强具身视觉语言模型在真实世界中的性能。数据集详情见https://dreamvu.ai/prism。
原文摘要 · Abstract (English)
A critical gap exists between the general-purpose visual understanding of state-of-the-art physical AI models and the specialized perceptual demands of structured real-world deployment environments. We present PRISM, a 270K-sample multi-view video supervised fine-tuning (SFT) corpus for embodied vision-language-models (VLMs) in real-world retail environments. PRISM is motivated by a simple observation - physical AI systems fail not because of poor visual recognition, but because they do not understand space, physical dynamics and embodied action well enough to operate reliably in the world. To this end, PRISM is grounded in a novel three-dimensional knowledge ontology that spans spatial knowledge, temporal and physical knowledge, and embodied action knowledge. It covers 20+ capability probes across four evaluation dimensions - Embodied Reasoning (ER), Common Sense (CS), Spatial Perception (SP), and Intuitive Physics (IP), and to our knowledge, PRISM is the first dataset to instantiate all three knowledge dimensions within a single real-world deployment domain. The corpus captures data from egocentric, exocentric and 360° viewpoints across five supermarket locations and includes open-ended, chain-of-thought, and multiple-choice supervision. At 4 fps, PRISM spans approximately 11.8M video frames and approximately 730M tokens, placing it among the largest domain-specific video SFT corpora. Fine-tuning on PRISM reduces the error rate across all 20+ probes by 66.6% over the pre-trained baseline, with significant gains in embodied action understanding where the accuracy improves by 36.4%. Our results suggest that ontology-structured, domain specific SFT can meaningfully strengthen embodied VLMs for real-world settings. The PRISM dataset and more details are available at https://dreamvu.ai/prism
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。