视觉语言模型助力第一人称视频理解,提升智能穿戴与机器人交互能力。
Vision-Language Models for Egocentric Video: From Hand-Object Interaction to Embodied AI

- 用图结构建模手-物-动作关系,捕捉时间动态
- 在长视频中识别交互意图准确率不足60%
- 适合研究可穿戴助手与具身智能的开发者
第一人称视频从佩戴者视角记录活动,直接反映注意力、手物交互与目标导向行为,对可穿戴智能、辅助系统、人机交互和具身智能至关重要。但其面临自我运动、遮挡、小物体、视角依赖外观及长时序依赖等挑战。视觉语言模型(VLMs)通过连接视觉与语义知识,结合自然语言监督,为解决这些问题提供了可能。本文综述了用于第一人称视频理解的VLMs,涵盖从传统识别架构到多模态基础模型与具身系统的演进。围绕任务、数据集、手物交互理解、时间推理、帧/片段选择、多模态表示学习、提示设计、语义对齐与模型适配展开分析。特别关注基于图与物体中心的推理机制,以建模随时间演变的手、物、动作与场景关系。进一步探讨第一人称感知与多模态基础模型如何支持可穿戴辅助、机器人技能学习、人向机器人技能迁移与具身决策。现有模型普遍更擅长识别可见物体,而对持续交互、动作与用户意图的识别仍不理想,尤其在长时间活动中。因此,亟需推进时间锚定推理、交互感知监督、高效长视频处理、多模态融合、图增强表示、跨域泛化、隐私保护与可信评估,以实现可部署的具身智能。
原文摘要 · Abstract (English)
Egocentric video captures activities from the wearer's perspective, providing a direct view of human attention, hand--object interaction, and goal-directed behavior. This perspective is increasingly important for wearable intelligence, assistive systems, human--robot interaction, and embodied AI, yet it introduces challenges including ego-motion, occlusion, small active objects, viewpoint-dependent appearance, and long-range temporal dependencies. Vision--language models (VLMs) offer a promising foundation for addressing these challenges by linking visual observations with semantic knowledge and natural-language supervision. This survey presents a critical review of VLMs for egocentric video understanding, tracing the progression from conventional recognition architectures to multimodal foundation models and embodied systems. We organize the literature around tasks, datasets, hand--object interaction understanding, temporal reasoning, frame and clip selection, multimodal representation learning, prompting, semantic alignment, and model adaptation. Particular attention is given to graph-based and object-centric reasoning as mechanisms for modeling relations among hands, objects, actions, and scene context over time. We further examine how first-person perception and multimodal foundation models support wearable assistance, robot skill learning, human-to-robot transfer, and embodied decision making. Across the reviewed literature, a consistent limitation emerges: current models recognize visible objects more reliably than evolving interactions, actions, and user intent, especially over long activities. We therefore identify temporally grounded reasoning, interaction-aware supervision, efficient long-video processing, multimodal fusion, graph-enhanced representations, cross-domain generalization, privacy, and trustworthy evaluation as key priorities for deployable embodied intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。