统一3D生成与理解,让模型既能造物又能识物。
UniMesh: Unifying 3D Mesh Understanding and Generation

- 用统一架构同时处理3D生成与理解任务
- 支持用户驱动的语义网格迭代编辑,效果显著提升
- 适合需要3D内容生成与分析一体化的研究者
近期3D视觉进展催生了专门用于3D理解(如形状分类、分割、重建)或3D生成(如合成、补全、编辑)的模型,但这些任务常被孤立处理,导致架构与表示碎片化,阻碍知识迁移和整体场景建模。为此,我们提出UniMesh,一个统一框架,在单一架构中联合学习3D生成与理解。首先,引入新型Mesh Head作为跨模型接口,连接基于扩散的图像生成与隐式形状解码器;其次,设计链式网格(Chain of Mesh, CoM),通过闭环的潜在空间、提示与重生成循环实现用户驱动的语义网格迭代编辑;第三,引入基于演员-评估器-自我反思三元组的自反思机制,诊断并修正高阶任务(如3D描述生成)中的失败。实验表明,UniMesh不仅在标准基准上表现优异,还实现了迭代编辑与生成-理解互增强等新能力。
原文摘要 · Abstract (English)
Recent advances in 3D vision have led to specialized models for either 3D understanding (e.g., shape classification, segmentation, reconstruction) or 3D generation (e.g., synthesis, completion, and editing). However, these tasks are often tackled in isolation, resulting in fragmented architectures and representations that hinder knowledge transfer and holistic scene modeling. To address these challenges, we propose UniMesh, a unified framework that jointly learns 3D generation and understanding within a single architecture. First, we introduce a novel Mesh Head that acts as a cross model interface, bridging diffusion based image generation with implicit shape decoders. Second, we develop Chain of Mesh (CoM), a geometric instantiation of iterative reasoning that enables user driven semantic mesh editing through a closed loop latent, prompting, and re generation cycle. Third, we incorporate a self reflection mechanism based on an Actor Evaluator Self reflection triad to diagnose and correct failures in high level tasks like 3D captioning. Experimental results demonstrate that UniMesh not only achieves competitive performance on standard benchmarks but also unlocks novel capabilities in iterative editing and mutual enhancement between generation and understanding. Code: https://github.com/AIGeeksGroup/UniMesh. Website: https://aigeeksgroup.github.io/UniMesh.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。