用一张图生成多样真实场景,解决长尾视觉数据稀缺问题
One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception

- 通过文本世界模型从单张图构建结构化场景,再扩展生成
- 合成数据训练的检测器接近真实数据性能,提升长尾场景识别
- 适合自动驾驶、海事监控等需要真实物理一致性的视觉任务
可靠的时空决策自动化(如自动驾驶、海事监控)依赖于稳健的视觉感知。然而,真实世界的时空数据存在严重异质性,关键安全场景常呈现极端长尾分布,导致数据稀疏与数据集偏移,降低检测性能并带来安全隐患。尽管合成数据生成是潜在解决方案,但现有扩散模型和生成对抗网络(GAN)通常缺乏显式空间对齐与结构约束,生成场景存在空间与物理不一致。为此,我们提出WMGen-v1,一种基于文本的世界模型框架,用于长尾空间数据生成。该框架利用大视觉语言模型(LVLM)从单张参考图像构建结构化场景表示,同时由大语言模型(LLM)在物理合理性与常识约束下进行引导式场景扩展。随后,基于此推理生成的结构化语义表示,扩散模型生成多样化且物理合理的长尾训练数据。在内部工业数据集ROADWork和LaRS基准上的实验表明,WMGen-v1优于基线方法。值得注意的是,仅使用WMGen-v1合成数据训练的检测器,在整体数据集指标上逼近仅使用真实数据的性能,凸显其缓解下游空间感知中长尾数据稀缺的潜力。
原文摘要 · Abstract (English)
Reliable spatial decision automation, such as autonomous driving and maritime surveillance, critically depends on robust visual perception. However, real-world spatiotemporal data exhibits severe heterogeneity, often manifesting as extreme long-tail distributions for safety-critical scenarios. This data scarcity induces dataset shift that degrades detection performance and pose safety risks. While synthetic data generation offers a potential solution, existing generative approaches, such as diffusion models and Generative Adversarial Networks (GANs), often lack explicit spatial grounding and structural constraints, resulting in spatial and physical inconsistencies in generated scenes. To address these challenges, we introduce WMGen-v1, an agentic text-based world model framework for long-tail spatial data generation. WMGen-v1 employs a Large Vision-Language Model (LVLM) to construct a structured scene representation from a single reference image, while a Large Language Model (LLM) performs guidance-based scene expansion under physical plausibility and commonsense constraints. Subsequently, conditioned on the structured semantic representations produced by this reasoning process, a diffusion model generates diverse and physically grounded long-tail training data. Experiments on internal industrial datasets, ROADWork, and LaRS benchmarks demonstrate that WMGen-v1 outperforms baseline approaches. Notably, detectors trained solely on WMGen-v1 synthetic data approach real-only performance on aggregate dataset-level metrics, highlighting its potential to alleviate long-tail data scarcity for downstream spatial perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。