arXiv:2603.07144cs.CV2026-03被引 1

构建32万件3D物体规范数据集,解决方向混乱问题。

CanoVerse: 3D Object Scalable Canonicalization and Dataset for Generation and Pose

  • 提出轻量级框架,秒级完成单个物体方向标准化
  • 数据集含1156类32万件物体,规模超前人一个数量级
  • 支持零样本点云朝向识别,适合3D生成与检索任务

3D学习系统隐含假设物体处于统一参考坐标系,但实际中每件资产都有任意全局旋转,模型需自行解决方向歧义。这种持续错位抑制了姿态一致的生成,并阻碍稳定方向语义的形成。为此,我们构建了 ameofmethod{},一个包含32万件物体、覆盖1,156个类别的大规模规范3D数据集——相比之前工作提升了一个数量级。在此规模下,方向语义可被统计学习:Canoverse提升了3D生成稳定性,实现了精确的跨模态3D形状检索,并使零样本点云朝向估计成为可能,即使面对分布外数据。这得益于一种新的规范化框架,通过紧凑假设生成与轻量级人工判别,将对齐时间从分钟级降至秒级,将规范化从人工整理转变为高吞吐数据生成流水线。数据集将在论文录用后公开。项目页:https://github.com/123321456-gif/Canoverse

原文摘要 · Abstract (English)

3D learning systems implicitly assume that objects occupy a coherent reference frame. Nonetheless, in practice, every asset arrives with an arbitrary global rotation, and models are left to resolve directional ambiguity on their own. This persistent misalignment suppresses pose-consistent generation, and blocks the emergence of stable directional semantics. To address this issue, we construct \methodName{}, a massive canonical 3D dataset of 320K objects over 1,156 categories -- an order-of-magnitude increase over prior work. At this scale, directional semantics become statistically learnable: Canoverse improves 3D generation stability, enables precise cross-modal 3D shape retrieval, and unlocks zero-shot point-cloud orientation estimation even for out-of-distribution data. This is achieved by a new canonicalization framework that reduces alignment from minutes to seconds per object via compact hypothesis generation and lightweight human discrimination, transforming canonicalization from manual curation into a high-throughput data generation pipeline. The Canoverse dataset will be publicly released upon acceptance. Project page: https://github.com/123321456-gif/Canoverse

3D生成数据集方向对齐规范表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。