arXiv:2607.10461cs.CVcs.AI2026-07

用几何自监督学习生成家具编码,无需标注即可表征类别与朝向。

Annotation-Free Furniture Codes: What They Encode, and How Far They Transfer

论文配图:Annotation-Free Furniture Codes: What They Encode, and How Far They Transfer
图 1 · 摘自论文原文
  • 通过无标签点云训练,用量化自动编码器提取家具几何特征。
  • 编码能准确恢复62.6%细分类别、85.6%大类和52.7度朝向信息。
  • 编码可迁移,但形状越有机越难跨数据集对齐,适合规则家具建模。

基于布局的3D场景合成器依赖人工标注的类别标签与标准姿态,本文探究仅用物体几何生成的单一自监督编码能否替代二者,并将编码作为独立表征进行研究。采用有限标量量化(FSQ)点云自编码器,在无标注、无姿态信息的放置式3D-FUTURE家具上进行Chamfer训练。诊断探测显示,仅从编码即可恢复细分类别(62.6±0.5%)、超类别(85.6±1.3%)及航向角(52.7±0.5°)。将Chamfer目标从旋转后点云切换为未旋转点云,导致航向信号消失但类别识别率上升,表明编码中的旋转信息由训练目标决定。跨资产库扩展需编码具备可迁移性;在未见数据集ShapeNet上,盒状家具可迁移,有机形态家具则不可,而目标无关增强部分弥补了差距。

原文摘要 · Abstract (English)

Layout-based 3D scene synthesizers place each object using two human-annotated channels: a categorical class label and a canonical-pose convention. We ask whether a single self-supervised token derived from object geometry can replace both, and study such tokens directly as a representation, decoupled from any synthesizer. A Finite Scalar Quantization (FSQ) point-cloud autoencoder is chamfer-trained on placed 3D-FUTURE furniture with no labels or pose annotations. Diagnostic probes recover fine-category (62.6 +/- 0.5%), super-category (85.6 +/- 1.3%), and yaw (52.7 +/- 0.5 deg) from the codes alone. Swapping the chamfer target from the rotated to the un-rotated point cloud collapses the yaw signal while raising class recovery, showing the codes' rotation content can be set by the training objective. Scaling across asset libraries needs codes that transfer; on an unseen dataset (ShapeNet), alignment is category-dependent: box-like furniture transfers, organically-shaped furniture does not, and a target-blind augmentation partly closes the gap.

自监督学习3D表征点云编码可迁移性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。