arXiv:2601.21449cs.ROcs.DC2026-01被引 2

统一合成数据生成框架,提升智能体训练的数据效率与稳定性。

Nimbus: A Unified Embodied Synthetic Data Generation Framework

  • 模块化四层架构,异步处理路径规划、渲染与存储。
  • 吞吐量提升2-3倍,支持大规模分布式长期运行。
  • 适合需要跨领域合成数据的智能体研究与工业级训练。

扩展数据量和多样性对提升智能体泛化能力至关重要。尽管合成数据生成可作为昂贵物理数据采集的可扩展替代方案,现有流水线仍呈碎片化且任务专用,导致工程效率低下和系统不稳定,难以满足基础模型训练所需的持续高吞吐数据生成需求。为此,我们提出Nimbus——一个统一的合成数据生成框架,旨在集成异构的导航与操作流水线。Nimbus采用解耦执行模型,将轨迹规划、渲染与存储分为异步阶段,通过动态调度、全局负载均衡、分布式容错及后端渲染优化,最大化利用CPU、GPU与I/O资源。评估表明,Nimbus相比未优化基线实现2-3倍的端到端吞吐提升,并在大规模分布式环境中保持稳健运行。该框架已作为InternData系列的数据生产核心,支持跨领域数据无缝合成。

原文摘要 · Abstract (English)

Scaling data volume and diversity is critical for generalizing embodied intelligence. While synthetic data generation offers a scalable alternative to expensive physical data acquisition, existing pipelines remain fragmented and task-specific. This isolation leads to significant engineering inefficiency and system instability, failing to support the sustained, high-throughput data generation required for foundation model training. To address these challenges, we present Nimbus, a unified synthetic data generation framework designed to integrate heterogeneous navigation and manipulation pipelines. Nimbus introduces a modular four-layer architecture featuring a decoupled execution model that separates trajectory planning, rendering, and storage into asynchronous stages. By implementing dynamic pipeline scheduling, global load balancing, distributed fault tolerance, and backend-specific rendering optimizations, the system maximizes resource utilization across CPU, GPU, and I/O resources. Our evaluation demonstrates that Nimbus achieves a 2-3X improvement in end-to-end throughput compared to unoptimized baselines and ensuring robust, long-term operation in large-scale distributed environments. This framework serves as the production backbone for the InternData suite, enabling seamless cross-domain data synthesis.

合成数据智能体分布式系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。