构建合成数据集GIFT,评估无纹理场景下的点跟踪性能
GIFT: Generated Indoor video frames for Texture-less point tracking
- 基于3D物体纹理强度分级,构建分层合成视频数据集
- 1800段室内视频,每段对应特定纹理等级,标注精准
- 首次系统评测点跟踪在无纹理场景下的表现差异
点跟踪在运动估计和视频编辑中日益成为关键方法,相比传统特征匹配,其在复杂相机轨迹和长时间序列下更具鲁棒性。然而,现有方法仍难以在无纹理或弱纹理区域稳定跟踪。本文首先提出用于评估3D物体纹理强度的指标,将ShapeNet中的3D模型分为三个纹理强度等级,并构建了GIFT——一个包含1800个室内视频序列的合成基准数据集,附带丰富标注。与以往数据集任意分配真实值不同,GIFT将真实值精确锚定在分类后的目标物体上,确保每段视频对应明确的纹理强度水平。我们在此基准上全面评估现有方法在不同纹理强度下的表现,并分析纹理对点跟踪的影响。
原文摘要 · Abstract (English)
Point tracking is becoming a powerful solver for motion estimation and video editing. Compared to classical feature matching, point tracking methods have the key advantage of robustly tracking points under complex camera motion trajectories and over extended periods. However, despite certain improvements in methodologies, current point tracking methods still struggle to track any position in video frames, especially in areas that are texture-less or weakly textured. In this work, we first introduce metrics for evaluating the texture intensity of a 3D object. Using these metrics, we classify the 3D models in ShapeNet into three levels of texture intensity and create GIFT, a challenging synthetic benchmark comprising 1800 indoor video sequences with rich annotations. Unlike existing datasets that assign ground truth points arbitrarily, GIFT precisely anchors ground truth on classified target objects, ensuring that each video corresponds to a specific texture intensity level. Furthermore, we comprehensively evaluate current methods on GIFT to assess their performance across different texture intensity levels and analyze the impact of texture on point tracking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。