arXiv:2504.00221cs.HCcs.AI2025-04被引 14

用眼动数据聚焦视频关键区域,大幅降低处理负荷仍保持高理解力。

GazeLLM: Multimodal LLMs incorporating Human Visual Attention

  • 基于眼动追踪定位视觉焦点,仅处理关键区域图像。
  • 仅需全分辨率1/10像素量,任务理解效果相当或更优。
  • 适合需要高效理解人类行为的智能助手与机器人应用。

大型语言模型正向多模态模型发展,可处理图像、音频、视频与文本。结合第一人称视角视频,多模态模型在理解人类活动方面展现出巨大潜力,可用于人机交互与人类增强应用,如活动辅助、现实世界代理及技能迁移至机器人。然而,高分辨率、长时视频产生大量潜在表示,带来显著内存与计算压力,限制了模型可处理的视频长度与分辨率。降低视频分辨率虽能减少内存占用,但常损害理解能力。本文提出一种结合眼动追踪数据优化第一人称视频分析的方法,将第一人称视觉视频分解为注视区域子块。通过仅处理这些被注视区域输入,该方法在任务理解上达到甚至优于全分辨率处理的效果,同时将输入视频像素数降低至十分之一,为多模态模型高效解析与利用人类技能提供了有效方案。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are advancing into Multimodal LLMs (MLLMs), capable of processing image, audio, and video as well as text. Combining first-person video, MLLMs show promising potential for understanding human activities through video and audio, enabling many human-computer interaction and human-augmentation applications such as human activity support, real-world agents, and skill transfer to robots or other individuals. However, handling high-resolution, long-duration videos generates large latent representations, leading to substantial memory and processing demands, limiting the length and resolution MLLMs can manage. Reducing video resolution can lower memory usage but often compromises comprehension. This paper introduces a method that optimizes first-person video analysis by integrating eye-tracking data, and proposes a method that decomposes first-person vision video into sub areas for regions of gaze focus. By processing these selectively gazed-focused inputs, our approach achieves task comprehension equivalent to or even better than processing the entire image at full resolution, but with significantly reduced video data input (reduce the number of pixels to one-tenth), offering an efficient solution for using MLLMs to interpret and utilize human skills.

多模态眼动追踪视觉聚焦效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。