arXiv:2605.10438cs.LGcs.CV2026-05

提出接口中心的生成状态,让3D模型可显式管理组件关系与装配逻辑。

Beyond Spatial Compression: Interface-Centric Generative States for Open-World 3D Structure

论文配图:Beyond Spatial Compression: Interface-Centric Generative States for Open-World 3D Structure
图 1 · 摘自论文原文
  • 用组件条件的规范局部令牌构建可查询的生成状态
  • 在多组件开放世界场景中零样本提升结构鲁棒性
  • 适合需要装配推理与结构修复的3D生成任务

当前3D分词器多将表示视为空间压缩:紧凑代码重建表面几何,但隐含组件归属与连接有效性。在存在交叉组件、噪声拓扑和弱规范结构的开放世界资产中,局部形状、组件身份与装配关系在潜在流中纠缠,解码时无法直接操作。本文提出接口中心的生成状态,将表示构建为可操作的状态而非被动压缩码。该状态显式暴露局部几何、组件归属与连接有效性作为可查询、约束与修复的变量。我们基于组件条件的规范局部令牌(C2LT-3D)实现此思想,将表示分解为规范局部几何、分区条件上下文与关系接缝变量。各因子分别应对压缩中心令牌的三类失效:姿态泄露、跨组件干扰与无效局部连接。该显式状态支持连接验证、潜在结构修复、定向干预与受约束序列化,无需额外后处理结构恢复模块。在单物体CAD模型上训练,零样本评估于开放世界多组件资产,结果表明其在对抗性连接设置下仍保持可操作性,显著提升结构鲁棒性。这提示开放世界3D生成表示应不仅以重建保真度衡量,更需评估其离散状态是否仍可用于装配级结构推理。

原文摘要 · Abstract (English)

Current 3D tokenizers largely treat representation as spatial compression: compact codes reconstruct surface geometry, but leave component ownership and attachment validity implicit. In open-world assets with intersecting components, noisy topology, and weak canonical structure, this creates a representation mismatch: local shape, component identity, and assembly relations become entangled in a latent stream and are not natively addressable during decoding. We formulate an alternative view, interface-centric generative states, in which tokenization constructs an operational state rather than a passive compressed code. The state exposes local geometry, component ownership, and attachment validity as variables that can be queried, constrained, and repaired during decoding. We instantiate this formulation with Component-Conditioned Canonical Local Tokens (C2LT-3D), factorizing representation into canonical local geometry, partition-conditioned context, and relational seam variables. Each factor targets a distinct failure mode of compression-centric tokens: pose leakage, cross-component interference, or invalid local attachment. This exposed state supports attachment validation, latent structural repair, targeted intervention, and constrained serialization without a separate post-hoc structure recovery module. Trained on single-object CAD models and evaluated zero-shot on open-world multi-component assets, C2LT-3D improves structural robustness and shows that its latent variables remain actionable under adversarial attachment settings. These results suggest that open-world 3D generative representations should be evaluated not only by reconstruction fidelity, but by whether their discrete states remain operational for assembly-level structural reasoning.

3D生成结构推理接口中心组件关系

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。