解决3D厨房数字孪生中的尺度模糊问题,实现语义与几何精准对齐。
KitchenTwin: Semantically and Geometrically Grounded 3D Kitchen Digital Twins
- 用视觉语言模型引导的几何锚点恢复真实尺度
- 通过重力对齐与曼哈顿结构约束提升几何一致性
- 适合需要高精度场景重建的具身智能研究者
具身AI的训练与评估需要以物体为中心、具有精确度量几何和语义定位的数字孪生环境。现有基于Transformer的前馈重建方法能从稀疏单目视频高效预测全局点云,但其存在固有的尺度模糊和坐标约定不一致问题,导致无法可靠融合这些无量纲点云与局部重建的物体网格。本文提出一种新的尺度感知3D融合框架,将视觉定位的物体网格与Transformer预测的全局点云注册,构建度量一致的数字孪生。方法引入视觉语言模型(VLM)引导的几何锚点机制,恢复真实世界度量尺度;并设计几何感知注册流程,通过重力对齐的垂直估计、曼哈顿世界结构约束及无碰撞局部优化,显式保证物理合理性。在真实室内厨房环境中实验表明,该方法显著提升跨网络物体对齐与几何一致性,支持多原语拟合与度量测量等下游任务。此外,我们公开了一个开源室内数字孪生数据集,包含度量标定场景以及语义标注且注册的物体中心网格注释。
原文摘要 · Abstract (English)
Embodied AI training and evaluation require object-centric digital twin environments with accurate metric geometry and semantic grounding. Recent transformer-based feedforward reconstruction methods can efficiently predict global point clouds from sparse monocular videos, yet these geometries suffer from inherent scale ambiguity and inconsistent coordinate conventions. This mismatch prevents the reliable fusion of these dimensionless point cloud predictions with locally reconstructed object meshes. We propose a novel scale-aware 3D fusion framework that registers visually grounded object meshes with transformer-predicted global point clouds to construct metrically consistent digital twins. Our method introduces a Vision-Language Model (VLM)-guided geometric anchor mechanism that resolves this fundamental coordinate mismatch by recovering an accurate real-world metric scale. To fuse these networks, we propose a geometry-aware registration pipeline that explicitly enforces physical plausibility through gravity-aligned vertical estimation, Manhattan-world structural constraints, and collision-free local refinement. Experiments on real indoor kitchen environments demonstrate improved cross-network object alignment and geometric consistency for downstream tasks, including multi-primitive fitting and metric measurement. We additionally introduce an open-source indoor digital twin dataset with metrically scaled scenes and semantically grounded and registered object-centric mesh annotations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。