arXiv:2509.15249cs.GRcs.AI2025-09

用因果推理生成更符合物理逻辑的3D场景。

Causal Reasoning Elicits Controllable 3D Scene Generation

  • 构建物体间的因果图,按逻辑顺序布局物体。
  • 通过因果干预确保空间配置符合物理规则。
  • 适合需要真实交互的虚拟场景设计者。

现有3D场景生成方法难以建模物体间的复杂逻辑依赖与物理约束,限制了其在动态、真实环境中的适应能力。本文提出CausalStruct框架,将因果推理融入3D场景生成。利用大语言模型(LLMs)构建因果图,节点代表物体及属性,边表示因果依赖与物理约束。通过因果顺序确定物体放置顺序,并应用因果干预调整空间配置以满足物理驱动约束,确保与文本描述和现实动态一致。优化过程中,使用比例-积分-微分(PID)控制器迭代调节物体尺度与位置。结合3D Gaussian Splatting与Score Distillation Sampling提升形状精度与渲染稳定性。大量实验表明,该方法生成的3D场景具备更强的逻辑一致性、真实的空间交互性与鲁棒适应性。

原文摘要 · Abstract (English)

Existing 3D scene generation methods often struggle to model the complex logical dependencies and physical constraints between objects, limiting their ability to adapt to dynamic and realistic environments. We propose CausalStruct, a novel framework that embeds causal reasoning into 3D scene generation. Utilizing large language models (LLMs), We construct causal graphs where nodes represent objects and attributes, while edges encode causal dependencies and physical constraints. CausalStruct iteratively refines the scene layout by enforcing causal order to determine the placement order of objects and applies causal intervention to adjust the spatial configuration according to physics-driven constraints, ensuring consistency with textual descriptions and real-world dynamics. The refined scene causal graph informs subsequent optimization steps, employing a Proportional-Integral-Derivative(PID) controller to iteratively tune object scales and positions. Our method uses text or images to guide object placement and layout in 3D scenes, with 3D Gaussian Splatting and Score Distillation Sampling improving shape accuracy and rendering stability. Extensive experiments show that CausalStruct generates 3D scenes with enhanced logical coherence, realistic spatial interactions, and robust adaptability.

3D生成因果推理场景建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。