让3D生成与理解更精准,通过动态对齐语义与几何层次。
ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation

- 用分层八叉树编码几何,引入锚点令牌实现跨模态精准匹配。
- 在图像到3D、文本到3D生成等任务上达到领先性能,计算量减半。
- 适合需要高效高精度3D生成与理解的研究者和开发者。
统一的3D基础模型旨在用单一主干网络完成3D资产生成与语言推理,但其文本-3D交互仍较隐式。现有方法将文本与3D token拼接为扁平序列,依赖自注意力机制,导致粗粒度结构线索与细粒度几何细节被混成单一表征。我们提出ELSA3D,通过弹性语义锚定机制,使语言与几何推理在对应抽象层级上协同进行。ELSA3D采用感知尺度的八叉树分词器表示几何,并引入锚点令牌(Anchor Tokens),作为稀疏跨模态单元,选择语义线索,将其路由至最相关的3D尺度,检索该尺度的几何证据,并将融合信号写回统一表征,保持交互稀疏而精确。轻量级块内路由器使计算与推理具有弹性,仅在语义对齐最需时激活锚点。ELSA3D在图像到3D生成、文本到3D生成及3D描述任务上均达当前最优,优于最强统一基线,同时相较非弹性版本减少约一半的浮点运算量(FLOPs)与推理延迟。
原文摘要 · Abstract (English)
Unified 3D foundation models aspire to generate 3D assets and reason about them in language within a single backbone, but their text-3D interaction remains largely implicit. Existing methods concatenate text and 3D tokens into a flat sequence and rely on self-attention, collapsing coarse structural cues and fine geometric details into one undifferentiated representation. We introduce ELSA3D, a unified 3D model that addresses this with elastic semantic anchoring, structuring language and geometric reasoning jointly along matched abstraction scales. ELSA3D represents geometry with a scale-aware octree tokenizer and introduces Anchor Tokens, sparse cross-modal units that select semantic cues, route them to the most relevant 3D scale, retrieve scale-specific geometric evidence, and write the fused signal back into the unified representation, keeping interaction sparse yet precise. A lightweight per-block router makes both computation and reasoning elastic, choosing which text tokens instantiate anchors at which geometric scale so that cross-modal capacity concentrates where alignment is most needed. ELSA3D achieves state-of-the-art performance across image-to-3D generation, text-to-3D generation, and 3D captioning, outperforming the strongest unified baseline while roughly halving FLOPs and inference latency relative to the non-elastic version of the same model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。