arXiv:2604.10125cs.CV2026-04被引 3

让3D室内场景生成更符合物理规律,提升机器人与AI应用可靠性。

PhyMix: Towards Physically Consistent Single-Image 3D Indoor Scene Generation with Implicit--Explicit Optimization

论文配图:PhyMix: Towards Physically Consistent Single-Image 3D Indoor Scene Generation with Implicit--Explicit Optimization
图 1 · 摘自论文原文
  • 用物理评估器量化几何、接触、稳定等9项约束,建立首个物理一致性基准
  • 通过训练与推理双阶段反馈,使生成场景在视觉与物理上均更合理
  • 适合需真实物理布局的机器人导航、智能设计等场景

现有单图像3D室内场景生成方法常生成视觉逼真但违背物理规律的结果,限制其在机器人、具身AI和设计领域的可靠性。为此,我们提出统一的物理评估器,衡量几何先验、接触关系、稳定性与可部署性四方面,进一步分解为九项子约束,构建首个物理一致性评估基准。基于此,分析显示当前先进方法仍严重缺乏物理意识。为此,我们提出PhyMix框架,将物理评估器反馈融入训练与推理过程,增强生成场景的物理合理性。具体包括:(i) 通过无评判器的场景组相对策略优化(Scene-GRPO)实现隐式对齐,利用评估器作为偏好信号,引导采样偏向物理可行布局;(ii) 通过即插即用的测试时优化器(TTO)进行显式修正,利用可微分评估信号在生成过程中纠正残余违规。整体方法统一了评估、奖励塑造与推理时修正,生成兼具视觉保真度与物理合理性的3D室内场景。大量合成评估验证了其在视觉保真度与物理一致性上的领先表现,风格化及真实图像的定性示例进一步展示了方法鲁棒性。代码与模型将在发表后公开。

原文摘要 · Abstract (English)

Existing single-image 3D indoor scene generators often produce results that look visually plausible but fail to obey real-world physics, limiting their reliability in robotics, embodied AI, and design. To examine this gap, we introduce a unified Physics Evaluator that measures four main aspects: geometric priors, contact, stability, and deployability, which are further decomposed into nine sub-constraints, establishing the first benchmark to measure physical consistency. Based on this evaluator, our analysis shows that state-of-the-art methods remain largely physics-unaware. To overcome this limitation, we further propose a framework that integrates feedback from the Physics Evaluator into both training and inference, enhancing the physical plausibility of generated scenes. Specifically, we propose PhyMix, which is composed of two complementary components: (i) implicit alignment via Scene-GRPO, a critic-free group-relative policy optimization that leverages the Physics Evaluator as a preference signal and biases sampling towards physically feasible layouts, and (ii) explicit refinement via a plug-and-play Test-Time Optimizer (TTO) that uses differentiable evaluator signals to correct residual violations during generation. Overall, our method unifies evaluation, reward shaping, and inference-time correction, producing 3D indoor scenes that are visually faithful and physically plausible. Extensive synthetic evaluations confirm state-of-the-art performance in both visual fidelity and physical plausibility, and extensive qualitative examples in stylized and real-world images further showcase the robustness of the method. We will release codes and models upon publication.

3D生成物理一致性场景生成测试时优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。