用运动对比三元组提升视觉语言模型时序一致性
ReMoT: Reinforcement Learning with Motion Contrast Triplets
- 构建大规模运动对比数据集,自动生成16.5K个三元组
- 新方法使时序推理性能提升25.1%,超越现有训练方式
- 适合关注机器人导航与自动驾驶的开发者
我们提出ReMoT,一种统一训练范式,系统解决视觉语言模型在时空一致性上的根本缺陷——这在导航、机器人和自动驾驶中尤为关键。ReMoT集成两个核心组件:(1) 基于规则的自动化框架,生成基于视频元标注的大规模(16.5K个三元组)运动对比数据集ReMoT-16K,优于人工或模型生成;(2) 群体相对策略优化,经实证验证在对比推理学习中表现最优且数据效率更高,远超标准监督微调。我们还构建首个细粒度运动对比三元组基准,用于评估视觉语言模型对细微运动属性(如相反方向)的区分能力。所获模型在新基准及多个标准视觉语言模型基准上达到顶尖水平,在时空推理任务上实现25.1%的显著性能提升。
原文摘要 · Abstract (English)
We present ReMoT, a unified training paradigm to systematically address the fundamental shortcomings of VLMs in spatio-temporal consistency -- a critical failure point in navigation, robotics, and autonomous driving. ReMoT integrates two core components: (1) A rule-based automatic framework that generates ReMoT-16K, a large-scale (16.5K triplets) motion-contrast dataset derived from video meta-annotations, surpassing costly manual or model-based generation. (2) Group Relative Policy Optimization, which we empirically validate yields optimal performance and data efficiency for learning this contrastive reasoning, far exceeding standard Supervised Fine-Tuning. We also construct the first benchmark for fine-grained motion contrast triplets to measure a VLM's discrimination of subtle motion attributes (e.g., opposing directions). The resulting model achieves state-of-the-art performance on our new benchmark and multiple standard VLM benchmarks, culminating in a remarkable 25.1% performance leap on spatio-temporal reasoning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。