用自然语言一键生成带元数据的定制化视频数据集
VDCook:DIY video data cook your MLLMs
- 通过自然语言指令和参数控制,自动完成视频检索与合成
- 支持多维度元数据标注,生成可追溯、可复现的数据包
- 适合需要快速构建垂直领域视频数据的研究团队
我们提出VDCook:一个自演化视频数据操作系统,为研究人员和垂直领域团队提供可配置的视频数据构建平台。用户通过自然语言查询和可调参数(规模、检索-合成比例、质量阈值)发起数据请求,系统自动执行查询优化,并行运行真实视频检索与受控合成模块。最终生成具有完整溯源信息和元数据的领域内数据包,以及可复现的Notebooks。与传统静态、一次性构建的数据集不同,VDCook基于MCP(Model Context Protocol)实现自动化数据摄入,使数据集成为动态演化的开放生态。系统还提供多维元数据标注(场景分割、运动评分、OCR比率、自动字幕等),为后续数据‘烹饪’与索引奠定基础。该平台旨在通过基础设施级解决方案大幅降低专业视频训练数据构建门槛,同时支持社区贡献与治理驱动的数据扩展范式。
原文摘要 · Abstract (English)
We introduce VDCook: a self-evolving video data operating system, a configurable video data construction platform for researchers and vertical domain teams. Users initiate data requests via natural language queries and adjustable parameters (scale, retrieval-synthesis ratio, quality threshold). The system automatically performs query optimization, concurrently running real video retrieval and controlled synthesis modules. It ultimately generates in-domain data packages with complete provenance and metadata, along with reproducible Notebooks. Unlike traditional static, one-time-built datasets, VDCook enables continuous updates and domain expansion through its automated data ingestion mechanism based on MCP (Model Context Protocol)\cite{mcp2024anthropic}, transforming datasets into dynamically evolving open ecosystems. The system also provides multi-dimensional metadata annotation (scene segmentation, motion scoring, OCR ratio, automatic captioning, etc.), laying the foundation for flexible subsequent data `cooking' and indexing\cite{vlogger}. This platform aims to significantly lower the barrier to constructing specialized video training datasets through infrastructure-level solutions, while supporting community contributions and a governance-enabled data expansion paradigm. \textbf{Project demo:} https://screenapp.io/app/v/WP0SvffgsH
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。