用文字控制生成风格多样的3D城市,还能精细编辑单个建筑。
MajutsuCity: Language-driven Aesthetic-adaptive City Generation with Controllable 3D Assets and Layouts
- 通过四阶段流程将城市拆解为可调控的布局、构件和材质组合。
- 相比现有方法,布局一致性提升83.7%,在多项指标上排名第一。
- 适合游戏、虚拟现实等领域需要高可控性3D城市生成的用户。
生成逼真3D城市是世界模型、虚拟现实与游戏开发的基础,理想的城市场景需兼具风格多样性、细节控制力与结构一致性。现有方法难以兼顾文本驱动的创意自由度与显式结构表示带来的对象级可编辑性。本文提出MajutsuCity,一个自然语言驱动且风格自适应的3D城市生成框架,将城市建模为可调控的布局、资产与材质组合,并采用四阶段流水线实现。为增强生成后的可编辑性,进一步引入MajutsuAgent——一个支持五类对象级操作的交互式语言引导编辑代理。同时构建MajutsuDataset,一个高质量多模态数据集,包含2D语义布局与高程图、多样化3D建筑资产、精选PBR材质与天空盒,均附带详细标注。我们还设计了一套涵盖结构一致性、场景复杂度、材质保真度与光照氛围的评估指标。大量实验表明,相较于CityDreamer,MajutsuCity降低83.7%布局FID;相比CityCraft,降低20.1%。在所有AQS与RDR评分中均排名第一,显著超越现有方法。结果验证了MajutsuCity在几何保真度、风格适应性与语义可控性方面的最先进水平。项目主页:https://longhz140516.github.io/MajutsuCity/
原文摘要 · Abstract (English)
Generating realistic 3D cities is fundamental to world models, virtual reality, and game development, where an ideal urban scene must satisfy both stylistic diversity, fine-grained, and controllability. However, existing methods struggle to balance the creative flexibility offered by text-based generation with the object-level editability enabled by explicit structural representations. We introduce MajutsuCity, a natural language-driven and aesthetically adaptive framework for synthesizing structurally consistent and stylistically diverse 3D urban scenes. MajutsuCity represents a city as a composition of controllable layouts, assets, and materials, and operates through a four-stage pipeline. To extend controllability beyond initial generation, we further integrate MajutsuAgent, an interactive language-grounded editing agent} that supports five object-level operations. To support photorealistic and customizable scene synthesis, we also construct MajutsuDataset, a high-quality multimodal dataset} containing 2D semantic layouts and height maps, diverse 3D building assets, and curated PBR materials and skyboxes, each accompanied by detailed annotations. Meanwhile, we develop a practical set of evaluation metrics, covering key dimensions such as structural consistency, scene complexity, material fidelity, and lighting atmosphere. Extensive experiments demonstrate MajutsuCity reduces layout FID by 83.7% compared with CityDreamer and by 20.1% over CityCraft. Our method ranks first across all AQS and RDR scores, outperforming existing methods by a clear margin. These results confirm MajutsuCity as a new state-of-the-art in geometric fidelity, stylistic adaptability, and semantic controllability for 3D city generation. We expect our framework can inspire new avenues of research in 3D city generation. Our project page: https://longhz140516.github.io/MajutsuCity/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。