arXiv:2604.09535cs.CV2026-04被引 2

构建真人自言自语链,提升长时序家务任务的视觉语言模型推理准确率。

EgoTL: Egocentric Think-Aloud Chains for Long-Horizon Tasks

论文配图:EgoTL: Egocentric Think-Aloud Chains for Long-Horizon Tasks
图 1 · 摘自论文原文
  • 通过说先于做记录每步目标与语音推理,实现时间对齐的思维链数据采集。
  • 在100+日常家务任务中验证,模型在长序列规划与空间定位上显著提升。
  • 适合研究视觉语言模型、具身智能及长时序任务建模的研究者使用。

大型基础模型在具身智能领域取得进展,可处理第一人称视角输入以完成家庭任务。然而,基于视觉语言模型(VLM)的自动标注常因缺乏准确的人类动作标签、思维链(CoT)和空间标注而产生噪声,这些错误在长时序空间指令遵循中被放大。问题源于对分钟级日常家务规划覆盖不足以及空间定位不准。导致模型推理链和世界模型生成出现幻觉物体、跳步或忽略真实物理属性。为此,我们提出EgoTL:一种第一人称数据的自言自语采集流水线。采用说先于做的协议,记录分步目标与带词级时间戳的口语推理;并通过度量尺度空间估计器、记忆库遍历和片段级标签,校准物理属性、场景上下文与导航/操作指令。利用EgoTL,我们在三个层级的六项任务维度上,对超过100个日常家务任务中的分钟级序列进行长时序生成与基准测试。结果表明,基础模型作为第一人称助手或开放世界模拟器仍存在明显不足。最后,我们使用与度量标签对齐的人类思维链,在EgoTL训练集上微调基础模型,显著改善了长时序规划、步骤推理、指令遵循与空间定位能力。

原文摘要 · Abstract (English)

Large foundation models have made significant advances in embodied intelligence, enabling synthesis and reasoning over egocentric input for household tasks. However, VLM-based auto-labeling is often noisy because the primary data sources lack accurate human action labels, chain-of-thought (CoT), and spatial annotations; these errors are amplified during long-horizon spatial instruction following. These issues stem from insufficient coverage of minute-long, daily household planning tasks and from inaccurate spatial grounding. As a result, VLM reasoning chains and world-model synthesis can hallucinate objects, skip steps, or fail to respect real-world physical attributes. To address these gaps, we introduce EgoTL. EgoTL builds a think-aloud capture pipeline for egocentric data. It uses a say-before-act protocol to record step-by-step goals and spoken reasoning with word-level timestamps, then calibrates physical properties with metric-scale spatial estimators, a memory-bank walkthrough for scene context, and clip-level tags for navigation instructions and detailed manipulation actions. With EgoTL, we are able to benchmark VLMs and World Models on six task dimensions from three layers and long-horizon generation over minute-long sequences across over 100 daily household tasks. We find that foundation models still fall short as egocentric assistants or open-world simulators. Finally, we finetune foundation models with human CoT aligned with metric labels on the training split of EgoTL, which improves long-horizon planning and reasoning, step-wise reasoning, instruction following, and spatial grounding.

具身智能思维链长时序任务第一人称数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。