开源280万条细粒度视频问答数据,构建可复现的视觉理解模型
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

- 采用无蒸馏训练流程,用大规模合成数据发现细节理解短板
- 发布280万条人工标注的视频问答对与时空定位描述
- 提供完整代码、数据与评估基准,支持透明科研
视觉-语言模型在计算机视觉研究中至关重要,但许多高性能模型仍为闭源,掩盖了其数据、设计与训练方法。研究社区常通过从黑箱模型中蒸馏来标注训练数据,虽取得良好基准表现,却阻碍了科学进步的可衡量性。本文提出在完全开放可复现的框架下构建感知语言模型(PLM),分析无需蒸馏的标准训练流程,并探索大规模合成数据以识别关键数据缺口,尤其在细粒度视频理解方面。为填补这些空白,我们发布了280万条人工标注的细粒度视频问答对及时空定位视频描述。同时引入PLM-VideoBench评估套件,用于评测涉及“什么”、“哪里”、“何时”和“如何”的挑战性视频理解任务。我们的工作通过提供数据、训练方案、代码与模型实现完全可复现,网址:https://github.com/facebookresearch/perception_models
原文摘要 · Abstract (English)
Vision-language models are integral to computer vision research, yet many high-performing models remain closed-source, obscuring their data, design and training recipe. The research community has responded by using distillation from black-box models to label training data, achieving strong benchmark results, at the cost of measurable scientific progress. However, without knowing the details of the teacher model and its data sources, scientific progress remains difficult to measure. In this paper, we study building a Perception Language Model (PLM) in a fully open and reproducible framework for transparent research in image and video understanding. We analyze standard training pipelines without distillation from proprietary models and explore large-scale synthetic data to identify critical data gaps, particularly in detailed video understanding. To bridge these gaps, we release 2.8M human-labeled instances of fine-grained video question-answer pairs and spatio-temporally grounded video captions. Additionally, we introduce PLM-VideoBench, a suite for evaluating challenging video understanding tasks focusing on the ability to reason about "what", "where", "when", and "how" of a video. We make our work fully reproducible by providing data, training recipes, code & models. https://github.com/facebookresearch/perception_models
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。