arXiv:2510.14992cs.CVcs.AI2025-10

用AI自动标注视频,让世界模型训练数据生成快80%还更安全。

GAZE:Governance-Aware pre-annotation for Zero-shot World Model Environments

  • 将360度视频转为标准视角并并行处理,实现高效预标注。
  • 通过多模态AI检测物体、语音、隐私内容,自动生成密集标签。
  • 减少人工审核量超80%,适合需高隐私合规的大规模数据集构建。

训练鲁棒的世界模型需要大规模、精确标注的多模态数据集,但传统人工标注耗时且昂贵。本文提出经生产验证的GAZE流水线,可将原始长视频自动转换为任务就绪的丰富监督信号。系统(i)将专有360度格式标准化为标准视角并切片以支持并行处理;(ii)集成场景理解、目标跟踪、语音转录、个人身份信息/不适宜内容/未成年人内容检测等AI模型,实现密集多模态预标注;(iii)将多源信号整合为结构化输出规范,便于快速人工校验。GAZE工作流显著提升效率(每小时审核节省约19分钟),并通过保守自动跳过低显著性片段,使人工审核量减少超过80%。通过提升标签密度与一致性,并融合隐私保护机制和链式保管元数据,该方法生成的高质量、隐私感知数据集可直接用于学习跨模态动态与动作条件预测。本文详细阐述了系统编排、模型选择与数据字典,提供了一种在不牺牲吞吐量或治理要求的前提下,规模化生成优质世界模型训练数据的可扩展范式。

原文摘要 · Abstract (English)

Training robust world models requires large-scale, precisely labeled multimodal datasets, a process historically bottlenecked by slow and expensive manual annotation. We present a production-tested GAZE pipeline that automates the conversion of raw, long-form video into rich, task-ready supervision for world-model training. Our system (i) normalizes proprietary 360-degree formats into standard views and shards them for parallel processing; (ii) applies a suite of AI models (scene understanding, object tracking, audio transcription, PII/NSFW/minor detection) for dense, multimodal pre-annotation; and (iii) consolidates signals into a structured output specification for rapid human validation. The GAZE workflow demonstrably yields efficiency gains (~19 minutes saved per review hour) and reduces human review volume by >80% through conservative auto-skipping of low-salience segments. By increasing label density and consistency while integrating privacy safeguards and chain-of-custody metadata, our method generates high-fidelity, privacy-aware datasets directly consumable for learning cross-modal dynamics and action-conditioned prediction. We detail our orchestration, model choices, and data dictionary to provide a scalable blueprint for generating high-quality world model training data without sacrificing throughput or governance.

世界模型自动标注多模态隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。