合成数据集SynthVerse提升点追踪泛化能力
SynthVerse: A Large-Scale Diverse Synthetic Dataset for Point Tracking
- 构建多样合成数据集,覆盖动画、交互操作等新场景
- 训练后追踪模型在跨域任务中准确率提升显著
- 适合需要强泛化能力的视觉追踪研究者使用
点追踪旨在复杂运动、遮挡和视角变化下持续跟踪视觉点,近年来随基础模型快速发展。但通用点追踪进展受限于高质量数据不足,现有数据集多样性有限且轨迹标注不完善。为此,我们提出SynthVerse,一个大规模、多样化的合成点追踪数据集。该数据集新增动画风格内容、具身操作、场景导航及关节物体等现有合成数据集中缺失的领域。SynthVerse通过覆盖更广物体类别并提供高质量动态运动与交互,显著提升数据多样性,支持更鲁棒的通用点追踪训练与评估。此外,我们建立了一个高度多样的点追踪基准,系统评估前沿方法在广泛领域漂移下的表现。大量实验表明,使用SynthVerse训练可一致提升泛化性能,并揭示现有追踪器在多样化设置下的局限性。
原文摘要 · Abstract (English)
Point tracking aims to follow visual points through complex motion, occlusion, and viewpoint changes, and has advanced rapidly with modern foundation models. Yet progress toward general point tracking remains constrained by limited high-quality data, as existing datasets often provide insufficient diversity and imperfect trajectory annotations. To this end, we introduce SynthVerse, a large-scale, diverse synthetic dataset specifically designed for point tracking. SynthVerse includes several new domains and object types missing from existing synthetic datasets, such as animated-film-style content, embodied manipulation, scene navigation, and articulated objects. SynthVerse substantially expands dataset diversity by covering a broader range of object categories and providing high-quality dynamic motions and interactions, enabling more robust training and evaluation for general point tracking. In addition, we establish a highly diverse point tracking benchmark to systematically evaluate state-of-the-art methods under broader domain shifts. Extensive experiments and analyses demonstrate that training with SynthVerse yields consistent improvements in generalization and reveal limitations of existing trackers under diverse settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。