arXiv:2603.19598cs.CV2026-03被引 2

用多模态图生成风格一致的室内场景,实现物体与整体风格协同控制。

FlowScene: Style-Consistent Indoor Scene Generation with Multimodal Graph Rectified Flow

  • 基于多模态图的三分支流模型,联合生成布局、形状与纹理。
  • 生成结果在真实度与风格一致性上优于基线方法。
  • 适合需要精细控制场景结构与外观的应用场景。

场景生成在工业领域有广泛应用,要求兼具高真实感和对几何与外观的精确控制。语言驱动的检索方法虽能从大型物体数据库中组合合理场景,但缺乏物体级控制且难以保证场景层面的风格一致性。基于图的方法可提升物体控制力并显式建模关系以保障整体一致性,但现有方法难以生成高保真纹理结果,限制了实际应用。我们提出 FlowScene,一种基于多模态图的三分支场景生成模型,协同生成场景布局、物体形状与纹理。核心是紧密耦合的修正流模型,在生成过程中交换物体信息,实现图上的协同推理。该机制支持对物体形状、纹理及关系的细粒度控制,同时确保结构与外观的场景级风格一致性。大量实验表明,FlowScene 在生成真实度、风格一致性以及人类偏好对齐方面均优于语言条件与图条件基线方法。

原文摘要 · Abstract (English)

Scene generation has extensive industrial applications, demanding both high realism and precise control over geometry and appearance. Language-driven retrieval methods compose plausible scenes from a large object database, but overlook object-level control and often fail to enforce scene-level style coherence. Graph-based formulations offer higher controllability over objects and inform holistic consistency by explicitly modeling relations, yet existing methods struggle to produce high-fidelity textured results, thereby limiting their practical utility. We present FlowScene, a tri-branch scene generative model conditioned on multimodal graphs that collaboratively generates scene layouts, object shapes, and object textures. At its core lies a tight-coupled rectified flow model that exchanges object information during generation, enabling collaborative reasoning across the graph. This enables fine-grained control of objects' shapes, textures, and relations while enforcing scene-level style coherence across structure and appearance. Extensive experiments show that FlowScene outperforms both language-conditioned and graph-conditioned baselines in terms of generation realism, style consistency, and alignment with human preferences.

场景生成图神经网络风格一致多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。