让3D生成能像人一样分步思考,精准理解语言指令
CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
- 用空间化思维表示分解3D潜在空间,实现几何分步推理
- 结合语言与空间推理,生成内容与描述高度一致
- 适合需要精准控制3D结构的生成任务
大型多模态模型的最新进展表明,明确的推理机制对提升模型可靠性、可解释性和跨模态对齐至关重要。尽管这类以推理为中心的方法在语言和视觉任务中已被证明有效,但其在3D领域的拓展仍不充分。CoRe3D提出一种统一的3D理解与生成推理框架,联合操作于语义和空间抽象层面,使来自语言的高层意图能直接指导底层3D内容的生成。该设计的核心是一种空间锚定的推理表示,将3D潜在空间分解为局部区域,使模型能够以组合式和过程化的方式推理几何结构。通过紧密耦合语义思维链推演与结构化空间推理,CoRe3D生成的3D输出展现出强局部一致性,并与语言描述高度对齐。
原文摘要 · Abstract (English)
Recent advances in large multimodal models suggest that explicit reasoning mechanisms play a critical role in improving model reliability, interpretability, and cross-modal alignment. While such reasoning-centric approaches have been proven effective in language and vision tasks, their extension to 3D remains underdeveloped. CoRe3D introduces a unified 3D understanding and generation reasoning framework that jointly operates over semantic and spatial abstractions, enabling high-level intent inferred from language to directly guide low-level 3D content formation. Central to this design is a spatially grounded reasoning representation that decomposes 3D latent space into localized regions, allowing the model to reason over geometry in a compositional and procedural manner. By tightly coupling semantic chain-of-thought inference with structured spatial reasoning, CoRe3D produces 3D outputs that exhibit strong local consistency and faithful alignment with linguistic descriptions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。