arXiv:2604.21575cs.CVcs.GR2026-04被引 1

无需已知尺度,一键适配多模态3D人体模型

OmniFit: Multi-modal 3D Body Fitting via Scale-agnostic Dense Landmark Prediction

论文配图:OmniFit: Multi-modal 3D Body Fitting via Scale-agnostic Dense Landmark Prediction
图 1 · 摘自论文原文
  • 用条件Transformer直接预测密集体表关键点,实现无尺度约束拟合
  • 在CAPE和4D-DRESS上达毫米级精度,性能提升57.1%~80.9%
  • 支持扫描、深度图、图像等多模态输入,适合真实与生成资产

将人体模型拟合到3D着装人体资产已被广泛研究,但多数方法仅依赖点云或多视角图像等单模态输入,且常需已知度量尺度,这在AI生成资产中难以实现,因尺度失真常见。我们提出OmniFit,一种可无缝处理全扫描、部分深度观测及图像捕获等多模态输入的方法,对真实与合成资产均保持尺度无关性。核心创新是采用简单有效的条件Transformer解码器,直接将表面点映射为密集体表关键点,进而用于SMPL-X参数拟合;另设可插拔图像适配器,利用视觉线索补足缺失几何信息。此外,引入专用尺度预测器,将受试者重缩放至标准人体比例。OmniFit在日常与宽松衣物场景下显著优于现有最佳方法,性能提升57.1%至80.9%。据我们所知,它是首个超越多视角优化基线的体形拟合方法,并首次在CAPE与4D-DRESS基准上实现毫米级精度。

原文摘要 · Abstract (English)

Fitting an underlying body model to 3D clothed human assets has been extensively studied, yet most approaches focus on either single-modal inputs such as point clouds or multi-view images alone, often requiring a known metric scale. This constraint is frequently impractical, especially for AI-generated assets where scale distortion is common. We propose OmniFit, a method that can seamlessly handle diverse multi-modal inputs, including full scans, partial depth observations, and image captures, while remaining scale-agnostic for both real and synthetic assets. Our key innovation is a simple yet effective conditional transformer decoder that directly maps surface points to dense body landmarks, which are then used for SMPL-X parameter fitting. In addition, an optional plug-and-play image adapter incorporates visual cues to compensate for missing geometric information. We further introduce a dedicated scale predictor that rescales subjects to canonical body proportions. OmniFit substantially outperforms state-of-the-art methods by 57.1 to 80.9 percent across daily and loose clothing scenarios. To the best of our knowledge, it is the first body fitting method to surpass multi-view optimization baselines and the first to achieve millimeter-level accuracy on the CAPE and 4D-DRESS benchmarks.

3D人体建模多模态融合尺度无关关键点预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。