让3D场景能预判互动变化,重建更真实可靠。
CA-World: Multi-Object Counterfactual Alignment for Efficient Interactive-Ready Reconstruction

- 用反事实对齐思想分离物体与背景,分步生成再整合
- 在多个测试中提升物体完整性和场景动态模拟效果
- 适合需要真实物理交互的机器人和自动驾驶研究
重建可交互的3D世界对物理仿真、虚拟现实、机器人及自动驾驶至关重要。现有方法多关注静态视觉保真度,缺乏对多物体交互的支持。本文认为,可交互重建应能预见潜在场景变化,保持几何完整性、视觉质量、多物体空间关系及物理合理性。受因果干预启发,提出CA-World框架,将反事实对齐学习融入解耦-重集成重建流程:将前景背景分离视为视觉干预,物体生成与背景补全视为反事实生成,场景重集成视为逆干预。根据反事实一致性,逆干预应回复原场景,从而建立重集成后与原始场景间的外观、空间与物理一致性对齐目标。该任务被形式化为有直接监督的反事实对齐学习。同时,利用物体级干预的局部性,通过观测场景约束反事实状态,实现高效一致的重集成,无需联合优化所有物体状态,降低计算开销与误差累积。在物体完整性、空间精度、室外背景补全、渲染质量、模拟动力学及下游应用上的实验验证了方法有效性。项目页面:https://chnxindong.github.io/ca-world/
原文摘要 · Abstract (English)
Reconstructing interaction-ready 3D worlds is essential for physical simulation, virtual reality, robotics, and autonomous driving. However, existing methods mainly optimize static and holistic visual fidelity, with limited support for multi-object interaction. We argue that an interaction-ready reconstruction should anticipate potential scene changes and preserve geometric completeness, visual quality, multi-object spatial relationship, and physical plausibility under potential interactions. To this end, motivated by the causal intervention, we propose CA-World, an efficient framework that integrates counterfactual alignment learning into a decoupling-reintegration reconstruction pipeline. Specifically, we formulate foreground-background decoupling as a visual intervention, separate object generation and background inpainting as counterfactual generation, and scene reintegration as an inverse intervention. According to counterfactual consistency, reversing the intervention should recover the factual world, motivating three alignment objectives between the reintegrated and original scenes: appearance, spatial, and physical consistency. This formulates interaction-ready reconstruction as counterfactual alignment learning with direct supervision. Moreover, leveraging the locality of object-level interventions, CA-World constrains counterfactual states using the observed scene, enabling efficient and coherent reintegration without jointly optimizing all object states, thereby reducing computational cost and error accumulation. Experiments on object completeness, spatial accuracy, outdoor background completion, rendering quality, simulated dynamics, and downstream applications demonstrate the effectiveness of CA-World. Project page: https://chnxindong.github.io/ca-world/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。