用课程学习提升视觉语言模型的动态推理能力
From Perception to Planning: Evolving Ego-Centric Task-Oriented Spatiotemporal Reasoning via Curriculum Learning

- 分三阶段构建课程式训练框架,逐步提升模型从感知到规划的能力
- 在长时序逻辑推理任务上达92.4%准确率,显著优于现有模型
- 适合研究具身智能、空间时序推理与机器人决策的开发者
现代视觉-语言模型在静态感知任务中表现优异,但在具身、第一人称任务所需的复杂时空推理方面仍受限。主要瓶颈在于其依赖被动视频数据学习的时间先验,常导致时空幻觉和动态环境中的泛化失败。为此,我们提出基于课程学习的EgoTSR框架,用于学习任务导向的时空推理。该框架基于一个核心假设:具身推理应从显式的空间理解,演进为内化的任务状态评估,最终实现长时程规划。为支持此范式,我们构建了大规模数据集EgoTSR-Data,包含4600万样本,分为三个阶段:思维链(CoT)监督、弱监督标签标注、长时程序列。大量实验表明,EgoTSR有效消除时间顺序偏差,在长时序逻辑推理任务上达到92.4%准确率,同时保持高精度细粒度感知能力,显著超越现有开源与闭源最先进模型。
原文摘要 · Abstract (English)
Modern vision-language models achieve strong performance in static perception, but remain limited in the complex spatiotemporal reasoning required for embodied, egocentric tasks. A major source of failure is their reliance on temporal priors learned from passive video data, which often leads to spatiotemporal hallucinations and poor generalization in dynamic environments. To address this, we present EgoTSR, a curriculum-based framework for learning task-oriented spatiotemporal reasoning. EgoTSR is built on the premise that embodied reasoning should evolve from explicit spatial understanding to internalized task-state assessment and finally to long-horizon planning. To support this paradigm, we construct EgoTSR-Data, a large-scale dataset comprising 46 million samples organized into three stages: Chain-of-Thought (CoT) supervision, weakly supervised tagging, and long-horizon sequences. Extensive experiments demonstrate that EgoTSR effectively eliminates chronological biases, achieving 92.4% accuracy on long-horizon logical reasoning tasks while maintaining high fine-grained perceptual precision, significantly outperforming existing open-source and closed-source state-of-the-art models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。