arXiv:2605.14399cs.CVcs.GR2026-05

通过3D场景干预生成一致的结构化监督信号,提升多模态学习性能。

SceneForge: Structured World Supervision from 3D Interventions

论文配图:SceneForge: Structured World Supervision from 3D Interventions
图 1 · 摘自论文原文
  • 基于可编辑3D世界,通过显式干预生成连贯监督信号。
  • 在多个基准上,物体与场景移除任务性能显著提升。
  • 适合需要干预一致性监督的视觉、生成与理解任务研究者。

许多多模态学习任务需要在编辑、视角变化和场景级干预下保持一致的监督信号,但观测数据集无法揭示底层场景状态及变化传播机制。本文提出SceneForge,一种基于干预驱动的框架,从可编辑的3D世界状态中生成结构化监督。该框架将每个场景表示为具有语义、几何与物理依赖关系的持久世界,通过施加显式干预(如物体移除或相机变化),并沿场景依赖关系传播其影响,生成与物体结构和场景效应一致的监督信号。由此得到的对偶观察、多视角输出及阴影、反射等效应感知信号,均源自共享世界状态,而非事后图像空间处理。我们基于Infinigen和Blender构建了无版权问题的室内监督资源,包含超过2000个场景,涵盖大量对偶样本与对齐标注,覆盖单视角与注册多视角设置。在相同训练预算下,引入SceneForge监督显著提升了多个基准上的物体移除与场景移除性能,定量与定性结果均验证其有效性。结果表明,将监督建模为可编辑世界的结构化状态转移,是实现干预一致性的可行且可扩展的多模态学习基础。

原文摘要 · Abstract (English)

Many multimodal learning tasks require supervision that remains consistent across edits, viewpoints, and scene-level interventions. However, such supervision is difficult to obtain from observation-level datasets, which do not expose the underlying scene state or how changes propagate through it. We present SceneForge, an intervention-driven framework that generates structured supervision from editable 3D world states. SceneForge represents each scene as a persistent world with semantic, geometric, and physical dependencies. By applying explicit interventions (e.g., object removal or camera variation) and propagating their effects through scene dependencies, SceneForge renders supervision that remains consistent with object structure and scene-level effects. This produces aligned outputs including counterfactual observations, multi-view observations, and effect-aware signals such as shadows and reflections, all derived from a shared world state rather than post hoc image-space processing. We instantiate SceneForge using Infinigen and Blender to construct a licensing-clean indoor supervision resource with a large number of counterfactual pairs and aligned annotations from over 2K scenes, covering both diverse single-view and registered multi-view settings. Under matched training budgets, incorporating SceneForge supervision improves both object removal and scene removal performance across multiple benchmarks in both quantitative and qualitative evaluation. These results indicate that modeling supervision as structured state transitions in editable worlds provides a practical and scalable foundation for intervention-consistent multimodal learning.

3D生成结构化监督多模态学习干预学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。