首个基于RGB的手持物体新视角合成基准,推动真实场景下视觉重建技术发展。
NVS-HO: A Benchmark for Novel View Synthesis of Handheld Objects
- 用手持和固定板双序列采集数据,实现真实环境下的多视角重建。
- 现有方法在自由手持条件下表现明显不足,存在显著性能差距。
- 适合研究新视角合成、3D重建及神经渲染的学者与工程师使用。
我们提出NVS-HO,首个针对真实环境中手持物体仅使用RGB输入的新视角合成基准。每个物体通过两种互补的RGB序列记录:(1) 手持序列,物体在静止相机前被操作;(2) 板面序列,物体固定在ChArUco板上,通过标记检测获得精确相机位姿。目标是让模型从(1)中学习物体完整外观,而(2)提供用于评估的真实图像。为建立基线,我们采用经典SfM流程和先进预训练前馈神经网络(VGGT)作为位姿估计器,并基于NeRF和高斯点云进行新视角合成建模。实验表明,在非约束手持条件下当前方法存在显著性能差距,凸显了对更鲁棒方法的需求。NVS-HO因此提供了一个具有挑战性的现实世界基准,推动基于RGB的手持物体新视角合成技术进步。
原文摘要 · Abstract (English)
We propose NVS-HO, the first benchmark designed for novel view synthesis of handheld objects in real-world environments using only RGB inputs. Each object is recorded in two complementary RGB sequences: (1) a handheld sequence, where the object is manipulated in front of a static camera, and (2) a board sequence, where the object is fixed on a ChArUco board to provide accurate camera poses via marker detection. The goal of NVS-HO is to learn a NVS model that captures the full appearance of an object from (1), whereas (2) provides the ground-truth images used for evaluation. To establish baselines, we consider both a classical SfM pipeline and a state-of-the-art pre-trained feed-forward neural network (VGGT) as pose estimators, and train NVS models based on NeRF and Gaussian Splatting. Our experiments reveal significant performance gaps in current methods under unconstrained handheld conditions, highlighting the need for more robust approaches. NVS-HO thus offers a challenging real-world benchmark to drive progress in RGB-based novel view synthesis of handheld objects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。