arXiv:2604.10385cs.CV2026-04被引 1

用游戏引擎生成带精确时空标注的视频数据,解决视频模型训练缺乏真实状态的问题。

GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models

论文配图:GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models
图 1 · 摘自论文原文
  • 基于事件图生成多角色同步视频,自动生成3D状态与空间关系等真值数据。
  • 数据集包含938段视频,空间关系密度是现有最佳水平的84倍,每帧都有完整实体关系图。
  • 适合研究视频理解、关系推理和模型评估的科研人员,尤其关注物理合理性与空间认知。

游戏引擎拥有视频模型难以学习的完整显式世界状态。我们将其转化为数据工具:GEST-Engine是一个生产级开源系统,可确定性地执行时空事件图(GESTs),无论是程序生成还是文本衍生,生成同步多角色场景视频,并在渲染时记录真值:3D实体与相机状态、成对空间关系、事件到帧映射、实例分割及长描述,边际标注成本为零。我们发布了GTASA,一个由该系统生成的938段视频样本,据我们所知,这是目前空间关系覆盖最密集的视频数据集:每帧均包含完整的实体对关系图,相较当前最优水平高约84倍。通过人类评估验证其物理合理性和语义一致性,前沿神经生成器在相同提示下表现不佳;定量上,使用GTASA预训练提升了视觉语言模型的视频描述性能。利用GTASA提供的精确3D真值,我们对六个冻结的视频编码器在11个时空任务中进行探查,首次实现了对冻结特征间实体间关系的精准探测,结果表明所有模型中‘谁靠近谁’的表现仅略高于随机水平。我们公开发布引擎、数据集与基准,使这一差距成为可测量、可训练的目标。

原文摘要 · Abstract (English)

Game engines hold what video models struggle to learn: a complete, explicit world state behind every frame. We turn one into a data instrument. GEST-Engine, our production-grade open-source system, deterministically executes Graphs of Events in Space and Time (GESTs), whether procedurally generated or derived from text, into videos of synchronized multi-actor scenarios, recording ground truth as it renders: 3D entity and camera state, pairwise spatial relations, event-to-frame mappings, instance segmentation, and long descriptions, at zero marginal annotation cost. With it we release GTASA, a 938-video sample of what the system can generate at arbitrary scale, carrying, to our knowledge, the densest spatial-relation coverage of any video dataset: a complete entity-pair relation graph at every frame, ~84x denser than the state of the art, frame-for-frame. We validate GTASA both qualitatively, through human evaluation of physical validity and semantic alignment where frontier neural generators, given the same prompts, largely fail, and quantitatively, with GTASA pretraining improving VLM video captioning. Probing six frozen video encoders across 11 spatio-temporal tasks enabled by GTASA's exact 3D ground truth, a previously untestable inter-entity relational probe of frozen video features, reveals that who-is-near-whom barely rises above chance for all of them. We release the engine, the corpus, and the benchmark, making this gap a measurable, trainable target.

视频生成时空分析真值数据关系推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。