arXiv:2608.01276cs.CV2026-08

用球面图引导扩散模型,实现无约束图像的全身捕捉。

Astrolabe: Spherical-Map Guidance Across Diffusion Pipelines for Full-Body Capture from Unconstrained Images

论文配图:Astrolabe: Spherical-Map Guidance Across Diffusion Pipelines for Full-Body Capture from Unconstrained Images
图 1 · 摘自论文原文
  • 通过固定球面映射生成空间噪声偏移,统一多阶段重建流程。
  • 在Puzzle-IOI和4D-Dress数据集上均提升图像与几何指标。
  • 无需密集配准或学习控制分支,适合通用图像重建系统。

从非受限照片中进行全身捕捉需在任意视角、姿态、裁剪和遮挡下建立全局对应关系。然而,该场景下的姿态、几何与基础特征估计过于不可靠,难以支持稠密匹配或外观迁移;而扩散修正器与优化流程缺乏统一接口以利用这些不确定对应。我们的洞察是:对应关系无需局部精确——其粗粒度视点与身体布局仍可组织扩散先验的适应与引导。我们提出 extit{Astrolabe},一种基于冻结视点引导球面图(SPH)的宿主可移植适配器。固定有界变换将 SPH 转换为空间噪声偏移,在先验适应阶段匹配,并在下游去噪或得分蒸馏引导中复用。当修正器暴露参考路由时,同一目标/参考 SPH 还提供粗略兼容性评分以选择原生外观特征;无路由优化仅使用共享偏移路径。因此,Astrolabe 采用单一 SPH--shift--adapt--guide 流程,避免稠密扭曲或学习控制分支。在 Puzzle-IOI 和 4D-Dress 上,它在所有主机及配对的 Puzzle-IOI 几何指标上均优于已有报告结果;图像增益扩展至背面视角,4D-Dress 几何整体保持稳定。

原文摘要 · Abstract (English)

Full-body capture from unconstrained photographs requires global correspondence across arbitrary views, poses, crops, and occlusions. Yet pose, geometry, and foundation features estimated in this setting are too unreliable for dense matching or appearance transfer, while diffusion rectifiers and optimization pipelines expose no common interface for consuming such uncertain correspondence. Our insight is that correspondence need not be locally accurate: its coarse viewpoint and body layout can still organize how a diffusion prior adapts and guides reconstruction. We introduce \emph{Astrolabe}, a host-portable adapter built on frozen viewpoint-guided spherical maps (SPH). A fixed bounded transform converts SPH into a spatial noise shift, which is matched during prior adaptation and reused during downstream denoising or score-distillation guidance in both pipeline categories. When a rectifier exposes a reference router, the same target/reference SPH additionally supplies coarse compatibility scores to select native appearance features; router-free optimization uses only the shared shift path. Astrolabe therefore follows one SPH--shift--adapt--guide process without dense warping or a learned control branch. Across Puzzle-IOI and 4D-Dress, it improves all reported image metrics in both hosts and all paired Puzzle-IOI geometry metrics; image gains extend to rear views, while 4D-Dress geometry remains stable overall.

全身捕捉扩散模型球面映射图像重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。