arXiv:2604.07882cs.CV2026-04被引 5

单视频快速重建物体外观与物理属性,无需人工标注

ReconPhys: Reconstruct Appearance and Physical Attributes from Single Video

  • 采用双分支自监督框架,直接从单视角视频预测几何、外观和物理属性
  • 未来预测PSNR达21.64,显著优于优化类方法的13.27,残差距离降至0.004
  • 推理时间小于1秒,适合机器人仿真与图形资产快速生成

非刚性物体的物理合理重建仍是重大挑战。现有方法依赖可微渲染进行逐场景优化,虽能恢复几何与动态,但需昂贵调参或人工标注,限制实用性与泛化能力。为此,我们提出ReconPhys,首个从前向框架中联合学习物理属性估计与3D高斯泼溅重建的方法,仅需单个单目视频输入。该方法采用双分支架构,通过自监督策略训练,无需真实物理标签。给定视频序列,ReconPhys可同时推断几何、外观与物理属性。在大规模合成数据集上的实验表明:本方法在未来预测上达到21.64 PSNR,远超最先进优化基线的13.27;切比雪夫距离从0.349降至0.004。关键优势在于推理速度极快(<1秒),相较现有方法数小时的耗时大幅提升效率,适用于机器人与图形学中的仿真资产快速生成。

原文摘要 · Abstract (English)

Reconstructing non-rigid objects with physical plausibility remains a significant challenge. Existing approaches leverage differentiable rendering for per-scene optimization, recovering geometry and dynamics but requiring expensive tuning or manual annotation, which limits practicality and generalizability. To address this, we propose ReconPhys, the first feedforward framework that jointly learns physical attribute estimation and 3D Gaussian Splatting reconstruction from a single monocular video. Our method employs a dual-branch architecture trained via a self-supervised strategy, eliminating the need for ground-truth physics labels. Given a video sequence, ReconPhys simultaneously infers geometry, appearance, and physical attributes. Experiments on a large-scale synthetic dataset demonstrate superior performance: our method achieves 21.64 PSNR in future prediction compared to 13.27 by state-of-the-art optimization baselines, while reducing Chamfer Distance from 0.349 to 0.004. Crucially, ReconPhys enables fast inference (<1 second) versus hours required by existing methods, facilitating rapid generation of simulation-ready assets for robotics and graphics.

三维重建物理模拟单视图高斯泼溅

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。