arXiv:2601.21199cs.CVcs.AI2026-01被引 2

Thinker模型提升机器人视觉语言理解,解决视角混淆与视频末尾信息遗漏问题。

Thinker: A vision-language foundation model for embodied intelligence

  • 构建包含第一人称视角和思维链数据的大型机器人感知数据集
  • 联合输入关键帧与完整视频,显著提升视频理解能力
  • 在任务规划基准上达到当前最优性能,适合机器人研发者参考

当大型视觉语言模型应用于机器人领域时,会面临人类轻易解决但模型易出错的问题,例如第三视角与第一视角混淆,以及在时间推理中忽略视频结尾信息。为此,我们提出Thinker,一种专为具身智能设计的大规模视觉语言基础模型。从两方面应对上述挑战:首先,构建一个大规模数据集,涵盖第一人称视频、视觉定位、空间理解及思维链数据;其次,提出一种简单有效的方法,通过同时输入关键帧与完整视频序列,显著增强模型对视频的理解能力。该模型在任务规划领域两个最常用基准数据集上取得当前最优结果。

原文摘要 · Abstract (English)

When large vision-language models are applied to the field of robotics, they encounter problems that are simple for humans yet error-prone for models. Such issues include confusion between third-person and first-person perspectives and a tendency to overlook information in video endings during temporal reasoning. To address these challenges, we propose Thinker, a large vision-language foundation model designed for embodied intelligence. We tackle the aforementioned issues from two perspectives. Firstly, we construct a large-scale dataset tailored for robotic perception and reasoning, encompassing ego-view videos, visual grounding, spatial understanding, and chain-of-thought data. Secondly, we introduce a simple yet effective approach that substantially enhances the model's capacity for video comprehension by jointly incorporating key frames and full video sequences as inputs. Our model achieves state-of-the-art results on two of the most commonly used benchmark datasets in the field of task planning.

具身智能视觉语言模型机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。